Intelligent track pavement damage segmentation method based on image instance segmentation

Through the improved YOLOV8s-D network, combined with the DDCNv4 module, BiLevel Routing Attention mechanism and DSConv convolution module, the problem of inability to detect magnetic nail damage, large environmental impact and poor adaptability in intelligent rail pavement damage detection is solved, and high-precision pavement damage segmentation and quantification is achieved.

CN120070888AActive Publication Date: 2025-05-30QINGDAO UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510130940.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The existing intelligent rail pavement damage detection method based on deep learning cannot detect magnetic nail damage, is greatly affected by the environment, and has poor adaptability to road surface defects of different scales.

Method used

A smart track road surface damage segmentation method based on image instance segmentation is proposed, and an improved YOLOV8s-D network is adopted. This network improves the recognition accuracy and adaptability of magnetic nail damage, marking line damage, cracks and potholes through the DDCNv4 module, BiLevel Routing Attention mechanism and DSConv convolution module.

Benefits of technology

It achieves an optimal segmentation accuracy of 92.7%, and can automatically identify and divide intelligent rail road damage. It is suitable for road defects of different lighting environments and different sizes, providing efficient damage quantification support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070888A_ABST
    Figure CN120070888A_ABST
Patent Text Reader

Abstract

A smart track pavement damage segmentation method based on image instance segmentation relates to the technical field of smart track pavement damage assessment, and comprises the following steps: step 1, making a data set; step 2, constructing a YOLOV8s-D network suitable for intelligent rail pavement damage segmentation based on the improvement of the YOLOV8s network, and performing hyper-parameter configuration; and step 3, GUI design. The YOLOV8s-D network suitable for intelligent rail pavement damage segmentation is provided, compared with other networks, magnetic nail damage, marker line damage, cracks and pits can be recognized, the recognition precision in different illumination environments is effectively improved, the YOLOV8s-D network can be suitable for pavement defects of different sizes, it is verified that the YOLOV8s-D network achieves the optimal segmentation precision of 92.7%, and the method is suitable for popularization and application. Automatic damage segmentation of the intelligent rail pavement is realized, and technical support is provided for subsequent damage quantification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent rail pavement damage assessment, and particularly relates to a method for segmenting intelligent rail pavement damage based on image instance segmentation. Background Art

[0002] Common asphalt pavement damages include cracks, potholes, and blurred markings. The damages on the intelligent rail pavement not only include the damages of common asphalt pavements but also its unique damage - magnetic nail damage. Existing deep learning image technologies for detecting common pavement damages have achieved a series of achievements and effects. However, the current methods for detecting intelligent rail pavement damages based on deep learning still have the following deficiencies: First, they cannot detect magnetic nail damages (missing nails, protrusions, depressions), and can only detect the characteristics of common asphalt pavement damages, such as cracks, potholes, and missing or blurred markings. Second, they are greatly affected by the environment. For example, the accuracy and precision of detection will be different under different lighting environments. Finally, they have poor adaptability to pavement defects of different scales and cannot effectively adapt to pavement defects of different sizes and shapes. Summary of the Invention

[0003] The present invention provides a method for segmenting intelligent rail pavement damage based on image instance segmentation, and proposes a YOLOV8s-D network suitable for intelligent rail pavement damage segmentation. Compared with other networks, it can identify magnetic nail damages, marking line damages, cracks, and potholes, effectively improve the recognition accuracy under different lighting environments, and can be applicable to pavement defects of different sizes. After verification, the present invention has achieved the best segmentation accuracy of 92.7%, realizing automatic damage segmentation of intelligent rail pavements and providing technical support for subsequent damage quantification.

[0004] To achieve the above object, the technical solution of the present invention is as follows: A method for segmenting intelligent rail pavement damage based on image instance segmentation, comprising the following steps: Step 1, dataset production: clarify the types and characteristics of 4 types of damages on the intelligent rail pavement, namely magnetic nail damage, marking line damage, cracks, and potholes, and construct a corresponding dataset for network training and verification; Step 2, improve the YOLOv8s network to construct a YOLOV8s-D network suitable for intelligent rail pavement damage segmentation, and perform hyperparameter configuration; Step 3, GUI design.

[0005] Preferably, step 1 includes the following specific steps: Obtain photos of the intelligent rail damage through automatic cruise shooting by a drone. All pictures are square. Use image augmentation technology to expand the data set. After expansion, divide the data set into a training set and a validation set. Label the intelligent rail damage through the labelme software to generate a json file. Draw the boundaries of the two types of damage with a closed curve to generate a mask file, and export a txt file. During the training process, the input image is resized to generate image samples with different width-to-height ratios.

[0006] Preferably, step 2 includes the following specific steps: Based on the YOLOv8s network architecture, design a DDCNv4 (Dilated Dense Convolutional Network v4) module to replace the C2F module in the backbone module, and add a BRA (BiLevel Routing Attention) attention mechanism. Add a DSConv (DepthwiseSeparable Convolution) convolution module in the Neck layer to replace the ordinary convolution in the Neck layer.

[0007] Preferably, in step 2, DDCNv4 refers to a convolutional neural network structure that combines dilated convolution and dense connection. The deformable convolution of DDCNv4 is used to dynamically adjust the position of the convolution kernel, so as to better adapt to intelligent rail road surface defects of different sizes and shapes; the dilated convolution is used to extract the contour information of large-size intelligent rail road surface defects to ensure the complete detection of the whole picture of intelligent rail road surface defects, thereby improving the adaptability of the model to detect intelligent rail road surface defects of different sizes.

[0008] Preferably, in step 2, aiming at the problem that the original YOLOv8s model has low recognition accuracy under the conditions of large differences in illumination conditions and partial occlusion of intelligent rail road surface defects, introduce the BiLevel Routing Attention attention mechanism, and simultaneously focus on the low-level and high-level features of the image through the BiLevel Routing Attention attention mechanism; the BiLevel Routing Attention attention mechanism is a dynamic, query-aware sparse attention mechanism, which is used to filter out most of the irrelevant key-value pairs at the rough region level so as to only retain a small part of the routing regions; secondly, apply fine-grained token-to-token attention in the union of these routing regions.

[0009] Preferably, in step 2, the BiLevel Routing Attention attention mechanism includes the following three parts: First, divide the feature map into S×S non-overlapping regions and perform a linear mapping: ; In the formula, Q: query, K: key, V: value, X r represents the region feature map after division, W q , W k , W v are the linear mapping weight matrices of query, key, and value respectively; Then, calculate the attention weights on the coarse-grained tokens, and only take the Topk regions as the relevant regions to participate in the fine-grained operations; ; ; In the formula, A r : Coarse-grained attention weight matrix, which represents the correlation between each coarse-grained region (or token); Q r : Coarse-grained query matrix, which is obtained from the original region feature map through linear mapping and is used to represent the information of other regions that the current region needs to pay attention to; K r : Coarse-grained key matrix, which is also obtained through linear mapping and is used to represent the information that other regions can be paid attention to; (K r ) T : The transpose matrix of K r , which is used for matrix multiplication with Q r to calculate the attention weights; I r : Index set of the Top-k regions, which contains the indexes of the top k most relevant regions for each query region in the coarse-grained attention weight matrix A r ; : A function used to select the region indexes corresponding to the top k largest weights from the attention weight matrix; Finally, use the Topk coarse-grained regions most relevant to each token as keys and values to participate in the final operation. To enhance locality, a depth convolution is still used on the keys and values: ; ; ; In the formula, K g : Fine-grained operation key matrix, which is obtained by aggregating the key vectors corresponding to the Top-k region indexes I r from the original key matrix K; V g : Fine-grained operation value matrix, which is obtained by aggregating the value vectors corresponding to the Top-k region indexes I r from the original value matrix V; : A function for aggregating corresponding vectors from the original matrix according to the index set; O: The final output feature map or output vector, which is obtained through fine-grained attention operation and Local Convolutional Enhancement (LCE). : The attention operation function is used to calculate the attention weights between the query vector and the aggregated key and value vectors, and perform weighted summation to obtain the output; Q: The original query matrix or the query vector after a certain transformation, which is used to calculate the attention weights with the aggregated key vector; LCE(V): The output of Local Convolutional Enhancement, which is obtained by applying depth convolution to the original value matrix V or the aggregated value matrix Vg to enhance the locality of the output.

[0010] Preferably, in step 2, when performing feature fusion of different dimensions, DSConv decomposes the traditional convolution kernel into two components: the Variable Quantization Kernel (VQK) and the distribution shift. It applies the kernel-based and channel-based distribution shifts to maintain the same output as the original convolution, and achieves lower memory usage by only storing integer values, aiming to improve the model speed and reduce the number of parameters.

[0011] Preferably, in step 2, during the quantization process of DSConv, the quantization function takes the number of bits to be quantized by the network as the input and saves the integer values using two's complement. For a number with bit length b, there is the following relationship: (1); In formula (1), represents the value of each parameter in the tensor; b: The quantization bit number, representing the number of bits used to represent the weight value; Z: The set of integers; N: The set of natural numbers; : The range of the quantized weight value, using two's complement to represent the integer value; First, scale the weights of each convolutional layer so that the maximum absolute value of the original weight matches the maximum value of the above quantization constraint; the new weight is stored in memory as an integer value for subsequent use in training and inference; by replacing ordinary convolution with DSConv, the memory saved for each tensor weight is: (2); In the formula, p: The proportion of memory saved for each tensor weight; b: The quantization bit number; 32: Usually represents the number of bits of a floating point number (such as float32); C i : The number of input channels; B: The block size or group size, which may be used for block processing of weights in quantization or grouped convolution; : Represents the number of blocks (rounded up or may vary according to the actual implementation) after dividing the number of input channels into blocks; By moving the VQK value through KDS and CDS, the weights of each block of the pre-trained network are stretched or rounded to fit the interval in Equation (1) and stored in VQK. The optimal value of KDS is: (3); In the formula, ξ: optimal scaling factor (or quantization scale factor); : candidate value of the scaling factor; w qi : quantized weight value; w i : original weight value; B: block size or number of samples, which represents the number of weights used to calculate the scaling factor here; Its closed form is: , (4).

[0012] Preferably, step 3 includes: the GUI interface includes 4 windows: image file selection, image display, recognition result, and save file.

[0013] Advantages of a method for intelligent rail pavement damage segmentation based on image instance segmentation according to the present invention: Based on the image instance segmentation method in deep learning, the present invention constructs an intelligent rail pavement damage image dataset for network training and verification; proposes a YOLOv8s-D network suitable for automatic segmentation of intelligent rail pavement damage to achieve pixel-level segmentation of cracks, potholes, marking line damage, and magnetic nail damage in a single image; this method is based on the YOLOv8s network architecture, replaces the C2F module in the backbone module with the DDCNv4 module, adds the BiLevel Routing Attention mechanism, and replaces the ordinary convolution in the Neck layer with the DSConv convolution module to improve the network segmentation effect and reduce the parameter calculation amount; in actual use, images can be collected through devices such as mobile phones, high-definition cameras, and drones, input into the GUI interface developed by the present invention for automatic damage segmentation, and the results are saved to provide data support for subsequent damage quantification by inspectors. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a data collection diagram taken by a drone according to the present invention.

[0015] Figure 2 is a data processing diagram of the present invention (in the figure: (a) labelme annotation diagram; (b) original image; (c) mask diagram; (d) label diagram).

[0016] Figure 3It is the YOLOv8s-D diagram of the intelligent rail pavement damage segmentation network architecture of the present invention (in the figure, Conv: Convolution; DDCNv4: A convolutional neural network structure combining dilated convolution and dense connection; BR-Attention: Double-layer routing attention mechanism, a neural network attention mechanism; DSConv: Distributed shift convolution; Concat: Concatenation; Upsample: Upsampling).

[0017] Figure 4 It is the framework diagram of the DDCNv4 module in the backbone network of the present invention (in the figure, conv: Convolution; dilation: Dilation).

[0018] Figure 5 It is the improved diagram of the attention mechanism of the present invention (in the figure, gather: Acquisition).

[0019] Figure 6 It is the DSConv network architecture of the present invention (in the figure, conv: Convolution).

[0020] Figure 7 It is the diagram of the training result of the present invention.

[0021] Figure 8 It is the comparison diagram of MIoU of different networks.

[0022] Figure 9 It is the comparison diagram of the segmentation effect diagrams of different networks.

[0023] Figure 10 It is the diagram of the ablation test result.

[0024] Figure 11 It is the segmentation effect diagram of the crack-bphdr dataset.

[0025] Figure 12 It is the segmentation effect diagram of the CrackForest dataset.

[0026] Figure 13 It is the GUI interface diagram of the intelligent rail pavement damage detection system. Detailed implementation manners

[0027] The following description details the implementation manners of the present invention in a step-by-step progressive manner. This description is only for the preferred embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

[0028] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, as well as a specific orientation structure and operation. Therefore, it should not be construed as a limitation to the present invention.

[0029] Embodiment 1: A method for intelligent rail pavement damage segmentation based on image instance segmentation according to the present invention, as Figures 1-13 shown, includes the following steps: Step 1, dataset production: clarify the types and characteristics of 4 types of damages, namely magnetic nail damage, marking line damage, cracks and potholes on the intelligent rail pavement, and construct a corresponding dataset for network training and verification; Step 2, improve the YOLOv8s network to construct the YOLOV8s-D network suitable for intelligent rail pavement damage segmentation, and perform hyperparameter configuration; Step 3, GUI design.

[0030] Embodiment 2: Based on Embodiment 1, this embodiment discloses that: the specific steps of the above-mentioned Step 1 include: the dataset is obtained by automatically cruising and shooting by a drone on the spot, as Figure 1 shown, use the DJI Air2 drone to automatically fly and shoot, and a total of 2000 intelligent rail damage photos are collected; all pictures are taken as squares. To avoid overfitting during model training, the dataset is expanded using image augmentation technology; after expansion, the number of images is 8 times that of the original dataset, and it is divided into a training set and a validation set according to a ratio of 7:3; then labelme software is used to label the intelligent rail damage to generate a json file, as Figure 2 shown in (a), draw the boundaries of the two damages with a closed curve to generate a mask file and export a txt file. The comparison of the original image, the marked image, the generated mask image and the visualized image is as Figure 2 shown in (b) to Figure 2 shown in (d); during the training process, in order to enhance the diversity of the data and improve the generalization ability of the model, the input images will be resized to generate image samples with different width-to-height ratios, and the ratios are set to (0.5, 0.75, 1.0, 1.25, 1.5 and 1.75).

[0031] Embodiment 3: Based on Embodiments 1 and 2, this embodiment discloses that: as Figure 3As shown, the step 2 includes the following specific steps: based on the YOLOv8s network architecture, a DDCNv4 (Dilated Dense Convolutional Network v4) module is designed to replace the C2F module in the backbone module, and a BRA (BiLevel Routing Attention) attention mechanism is added, and a DSConv (Depthwise Separable Convolution) convolution module is added to the Neck layer to replace the ordinary convolution of the Neck layer.

[0032] Embodiment 4: Based on Example 3, this embodiment discloses: Figure 4 As shown, in the step 2, DDCNv4 refers to a convolutional neural network structure that combines dilated convolution and dense connection. Aiming at the problem that the original semantic segmentation model of YOLOv8 has poor adaptability to smart rail pavement defects of different scales, a DDCNv4 module is designed to replace the C2F module in the backbone module; the deformable convolution of DDCNv4 can dynamically adjust the position of the convolution kernel to better adapt to smart rail pavement defects of different sizes and shapes; because the receptive field of ordinary convolution is limited, in order to obtain the global information of large-sized smart rail pavement defects, dilated convolution is introduced; through the above-mentioned dilated convolution, the contour information of large-sized smart rail pavement defects can be better extracted to ensure the complete detection of the overall picture of smart rail pavement defects, thereby improving the adaptability of the model to detect smart rail pavement defects of different sizes.

[0033] Embodiment 5: Based on Example 4, this embodiment discloses: in the step 2, the attention mechanism is a neural network resource allocation scheme for allocating computing resources to more important tasks; the existing attention module is usually inherited into each block to improve the output from the previous layer; in order to solve the problem that the original model of YOLOv8s has low recognition accuracy due to large differences in detection lighting conditions and partial occlusion of smart rail pavement defects, the BiLevel Routing Attention mechanism is introduced; the BiLevel Routing Attention mechanism can pay attention to both low-level and high-level features of the image; low-level features (such as edges, textures, etc.) help to identify detailed information of asphalt surface defects, while high-level features (such as shape, structure, etc.) help to understand the overall outline and type of smart rail pavement defects.

[0034] Embodiment 6: Based on Example 5, this embodiment discloses: Figure 5As shown, in the said step 2, the BiLevelRouting Attention mechanism includes the following three parts: First, divide the feature map into S×S non-overlapping regions and perform a linear mapping: ; Then, calculate the attention weights on the coarse-grained tokens, and only take the Topk regions among them as relevant regions to participate in the fine-grained operations; ; ; Finally, take the Topk coarse-grained regions most relevant to each token as keys and values to participate in the final operation. To enhance locality, a depth convolution is still used on the keys and values: ; ; .

[0035] Example 7: Based on Example 6, this example discloses that: in the said step 2, the designed DDCNv4 module can better solve the problem of difficult recognition of different scales, and introducing the BiLevel Routing Attention mechanism can solve the problem of high difficulty in recognizing intelligent rail pavement defects and complex backgrounds; however, the DDCNv4 module introduces the learning and calculation of offsets, so it will increase the complexity and computational overhead of model inference, which increases the computational complexity and inference time of the model and is not suitable for real-time monitoring of multiple defects on the intelligent rail pavement; secondly, the calculation and storage of the feature matrix offset require additional memory, which will increase the demand for hardware resources; in order to reduce the number of model parameters as much as possible without affecting the model accuracy, the DSConv convolution module is used to replace the ordinary convolution in the Neck layer; when performing feature fusion of different dimensions, the DSConv convolution decomposes the traditional convolution kernel into 2 components: a variable quantization kernel (VQK) and a distribution offset, and applies kernel-based and channel-based distribution offsets to maintain the same output as the original convolution; by only storing integer values, lower memory usage is achieved, and the purpose of improving the model speed and reducing the number of parameters is achieved. The network structure of DSConv is as Figure 6 shown.

[0036] Example 8: Based on Example 7, this example discloses that: In the said step 2, during the quantization process of DSConv, the quantization function takes the number of bits to be quantized by the network as input and uses two's complement to save integer values. For a number with a bit length of b, there is the following relationship: (1); In formula (1), represents the value of each parameter in the tensor; First, scale the weights of each convolutional layer so that the maximum absolute value of the original weights matches the maximum value of the above quantization constraint; the new weights are stored in memory as integer values for subsequent use in training and inference; by replacing ordinary convolution with DSConv, the memory saved for each tensor weight is: (2); By using KDS and CDS to shift the VQK values, the weights of each block of the pre-trained network are stretched or rounded to fit the interval in formula (1) and stored in VQK, where the optimal value of KDS is: (3); Its closed form is: , (4).

[0037] As described above, by choosing to replace the ordinary convolution in the Neck layer of YOLOv8 with DSConv convolution, the purpose of accelerating the model training speed and reducing the model storage space is achieved.

[0038] Example 9: Based on Example 8, this example discloses that step 3 includes: establishing a GUI interface for intelligent rail pavement damage segmentation to achieve the automatic segmentation effect of damaged targets on the intelligent rail pavement. The GUI interface includes 4 windows: image file selection, image display, recognition result, and save file; the main GUI interface is as Figure 13 shown: Click "Upload Image" to select the image to be segmented, which will be presented in the upper left part of the system. Subsequently, the system automatically segments the damaged area in the image, and the image after network segmentation is obtained in the upper right part of the system. Finally, click "Save Image" to save the file with a custom path. This system helps the detection personnel make a quick judgment on the severely damaged area of the intelligent rail pavement, take corresponding measures, and provide necessary technical support for subsequent damage quantification.

[0039] Example 10: Based on Example 9, this example discloses the implementation part of network training and performance evaluation, which is as follows: 10.1 Test environment: The test environment for damage segmentation based on YOLOv8s-D is shown in Table 1: Table 1 Test environment configuration table: ; 10.2 Evaluation metrics: To quantify the performance of the model, pixel accuracy (P.A.) and mean intersection over union (MIoU) are set. The former is the proportion of correctly predicted positive samples among all correctly predicted samples, and the latter is the ratio of the intersection to the union of the ground truth and the predicted value; the calculation method of P.A. is shown in Equation (5); the IoU expression is shown in Equations (6) and (7), and the average value of the ratio of the intersection over union of all classes is calculated as MIoU, as shown in Equation (8): (5); (6); (7); (8); In Equation (5), the numerator is the sum of the number of all correctly predicted pixels, and the denominator is the total number of pixels in the image. FN (False Negative), FP (False Positive), TN (True Negative), and TP (True Positive) are the values in the confusion matrix; during network training, the loss function is used as an indicator to measure the deviation between the predicted value and the ground truth. In this embodiment, CIou (Complete IoU) is selected as the loss function. Based on DIoU (Distance-IoU), it can also consider the aspect ratio of the two rectangles simultaneously, that is, the similarity of shapes. The original intention of the CIou design is to introduce an aspect ratio influence factor to make up for the deficiency of DIoU in expressing the distance between the bounding boxes, enabling the model to more accurately distinguish cracks and bursts. By adopting CIou, the model can better learn the characteristics of cracks and bursts, thereby improving the accuracy of segmentation; 10.3 Training Results: The parameters of the backbone network layer of the model are initialized with the weights pre-trained on the ImageNet dataset, which is crucial for improving the model performance and accelerating the convergence speed. It is set to save a checkpoint (CheckPoint) every 50 iterations to save the model weights, and record and output a log message each iteration to track the training status of the model in real time; in terms of optimization, the Adam optimizer is selected to adjust the model parameters to minimize the loss function, and the gradient overflow problem is handled by the dynamic loss scaling technique. The momentum parameter of the Adam optimizer is set to (β 1 = 0.937, β 2(=0.999), the learning rate was initialized to 0.01; to prevent the model from overfitting and reduce its complexity, a weight decay strategy was adopted, and the decay coefficient was set to 0.0005; the adjustment of the learning rate started from the initial stage of training and was dynamically adjusted according to the number of iterations, and the adjustment stopped after reaching 500 iterations; the initial learning rate was set to 0.01; a total of 500 iterations were performed during the entire training process, and validation was performed every 20 iteration cycles during this period; Figure 7 (a - c) The comprehensive display of the training and validation results includes a series of metrics such as box loss, classifyloss - CLS loss, segmentation loss, DFL loss, precision, recall, and mean average precision (mAP). "B" and "M" represent boxes and masks. A careful examination of these results reveals the convergence trend of all loss metrics, precision, and mAP as the number of iterations increases; from the data results, the loss values gradually decrease and tend to be stable. Especially in the later stage of the training process, the loss values hardly change. The convergence of the training process indicates that the model effectively learns to minimize the error and improves the prediction accuracy over time; mAP50 represents the mean precision calculated at an IoU threshold of 0.5; as can be seen from Figure 7 it, metrics / mAP50 (Box) is close to 0.9; at a more stringent 50 - 95% IoU (mAP50 - 95), it also shows a performance higher than 0.8, further verifying that the model has strong segmentation performance and robustness under strict evaluation criteria; in summary, the final loss value is low and stable, achieving a good balance between precision and recall, and the model performs excellently when segmenting damage defect targets at different IoU thresholds, effectively learning the defect features without overfitting, and achieving good training results.

[0040] Example 11: Based on Example 10, this example discloses the implementation part of the comparative experiment, which is specifically as follows: To further demonstrate the excellent performance of the YOLOv8s - D model, multiple state - of - the - art models such as YOLOv8s, UNet, DeepLabV3, and LR - ASPP were used to train and test the same dataset. By comparing the MIoU of each model, the performance of each network was evaluated targeted. The training results are as Figure 8 shown; To more intuitively display the model segmentation effect and verify the accuracy of the model for the segmentation of intelligent rail pavement damage, Figure 9 is the segmentation comparison result of 5 networks after training; From Figure 8 and Figure 9 it can be clearly seen that: (1) The detection and segmentation effects of the 5 networks on the marking lines and pits are good, but there are significant differences in the segmentation of magnetic nail damage and cracks. The YOLOv8s-D network has the best effect, and the final MIoU can reach 93.5%, which can clearly and accurately detect and segment each crack and each magnetic nail damage. The LR-ASPP network has the worst effect, and the final MIoU can only reach 40.2% and will no longer increase. The segmentation integrity of the target object is insufficient, and the detection error rate is relatively high. When multiple damages such as missing nails, pits, and cracks coexist, the YOLOv8s-D network can accurately classify and segment the damages, but for other networks, effective classification and segmentation cannot be performed, which reflects the advantage of the instance segmentation network compared to the semantic segmentation network; (2) Compared with the unimproved network YOLOv8s, the MIoU of YOLOv8s-D has increased by 3.7%, the number of detected cracks and magnetic nail damages is larger and more accurate, and the segmentation accuracy is higher; In summary, the YOLOv8s-D network proposed by the present invention is superior to other networks in terms of the overall performance, edge description, and detail capture of images. By improving the original network structure and making up for its defects, the YOLOv8s-D network successfully improves the segmentation accuracy of target information and demonstrates strong robustness and stability.

[0041] Example 12: Based on Example 11, this example discloses the implementation part of the ablation experiment, which is specifically as follows: The purpose of the ablation experiment is: to construct an ablation comparison experiment to verify the effectiveness of the optimization technology, ensure that there is no competition and conflict between various methods, optimize resource utilization, and avoid inconsistencies and chaos in the model training and decision-making processes; the present invention uses ablation experiments on the DDCNv4, BRA, and DSConv modules to confirm the effectiveness of the enhancement; in order to evaluate how these changes affect the performance, YOLOv8s is used as a benchmark for ablation; Table 2 Data table of ablation experiment: ; From Table 2 and Figure 10As shown in the figure, the YOLOv8s algorithm has added DDCNv4, BRA, and DSConv modules, and has been improved in terms of three indicators: the number of parameters, mAP0.5, and FPS. Compared with the YOLOv8s algorithm, the YOLOv8s-D network has increased by 5.9% in mAP0.5. After adding DSConv, the FPS (frames Per Second) has increased significantly by 19.72. After adding the DDCNv4 module and the BRA attention mechanism, although the FPS has decreased, the mAP0.5 has increased significantly by 4.3%. The experimental data shows that each module of the proposed method has an improvement effect on the model, and the combined use has a better effect than the single use, indicating that the new model is superior to the original YOLOv8s model in the target segmentation task.

[0042] Example 13: Based on Example 12, this example discloses the implementation part of the dataset experiment, which is specifically as follows: In order to verify the generalization ability of the network provided by the present invention, multiple publicly available crack datasets, including the crack-bphdr dataset and CrackForest, were used to comprehensively test the network proposed by the present invention. Among them, the crack-bphdr dataset includes about 4000 concrete crack images under different backgrounds, which were taken under different lighting conditions, angles, and crack types to ensure the diversity and generalization ability of the dataset. The segmentation results are as Figure 11 shown; CrackForest includes 712 asphalt pavement crack images taken under various lighting and photographic conditions, including single cracks and intersecting cracks. The segmentation results using this dataset are as Figure 12 shown; (1) It can be clearly seen from Figure 11 that the YOLOv8s-D network has excellent segmentation performance, and the predicted MIoU value is between 0.87 and 0.95, with excellent results. This is because the characteristics and background information of the cracks in the crack-bphdr dataset are relatively obvious, making the segmentation easier; (2) It can be known from Figure 12 that when using the YOLOv8s-D network to segment single cracks, the MIoU value is generally between 0.85 and 0.95, showing good segmentation results; when segmenting intersecting cracks and multiple cracks, the MIoU value is generally between 0.70 and 0.95, and the segmentation accuracy has decreased compared to single crack segmentation. This is because the crack morphology in this dataset is more complex, increasing the difficulty for the model to segment the cracks from the background and resulting in a decrease in segmentation efficiency; The test results of integrating two publicly available crack datasets show that the YOLOv8s-D network has excellent performance and strong generalization ability in the crack segmentation task. This network can accurately identify and segment cracks of various shapes, scales, and complexities, showing not only excellent performance on specific datasets but also stable and efficient performance in cross-dataset tests.

Claims

1. A smart track pavement damage segmentation method based on image instance segmentation, characterized by: The steps include: Step 1: Dataset creation: Identify the types and characteristics of four types of damage, namely magnetic nail damage, marking line damage, cracks and potholes on the smart track pavement, and construct corresponding datasets for network training and verification; Step 2: Based on the improvement of the YOLOv8s network, a YOLOV8s-D network suitable for intelligent rail pavement damage segmentation is constructed, and hyperparameters are configured; Step 3: GUI design.

2. The intelligent track pavement damage segmentation method based on image instance segmentation as claimed in claim 1, characterized in that: The step 1 includes the following specific steps: obtaining photos of smart rail damage through automatic cruising of a drone, all pictures are square, expanding the data set using image augmentation technology, and dividing the data set into a training set and a validation set after expansion, annotating the smart rail damage through labelme software, generating a json file, drawing the boundaries of the two types of damage with closed curves, generating a mask file, and exporting a txt file. During the training process, the input image is resized to generate image samples with different aspect ratios.

3. The intelligent track pavement damage segmentation method based on image instance segmentation as claimed in claim 2, characterized in that: The step 2 includes the following specific steps: based on the YOLOv8s network architecture, a DDCNv4 module is designed to replace the C2F module in the backbone module, and an attention mechanism is added, and a DSConv convolution module is added to the Neck layer to replace the ordinary convolution of the Neck layer.

4. The intelligent track pavement damage segmentation method based on image instance segmentation as claimed in claim 3 is characterized by: In the step 2, DDCNv4 refers to a convolutional neural network structure that combines dilated convolution and dense connections. The deformable convolution of DDCNv4 is used to dynamically adjust the position of the convolution kernel to better adapt to smart rail pavement defects of different sizes and shapes; the dilated convolution is used to extract the contour information of large-size smart rail pavement defects to ensure the complete detection of the entire picture of smart rail pavement defects, thereby improving the adaptability of the model to detect smart rail pavement defects of different sizes.

5. The method for intelligent track pavement damage segmentation based on image instance segmentation as claimed in claim 4, characterized in that: In the step 2, in order to solve the problem that the original YOLOv8s model has low recognition accuracy when detecting large differences in lighting conditions and when defects in the smart rail pavement are partially obscured, an attention mechanism is introduced. The attention mechanism simultaneously focuses on low-level and high-level features of the image; the attention mechanism is a dynamic, query-aware sparse attention mechanism, which is used to filter out most of the irrelevant key-value pairs at the coarse area level so that only a small part of the routing area is retained; secondly, fine-grained token-to-token attention is applied to the union of these routing areas.

6. The method for intelligent track pavement damage segmentation based on image instance segmentation as claimed in claim 5, characterized in that: In step 2, the attention mechanism includes the following three parts: First, the feature map is divided into S×S non-overlapping regions and linearly mapped: ; Then, the attention weight is calculated on the coarse-grained Token, and only the Topk regions are taken as relevant regions to participate in the fine-grained calculation; ; ; Finally, the most relevant Topk coarse-grained regions of each token are used as keys and values ​​for the final operation. In order to enhance locality, a deep convolution is used on the keys and values: ; ; 。 7. The method for intelligent track pavement damage segmentation based on image instance segmentation as claimed in claim 6, characterized in that: In step 2, when fusing features of different dimensions, the DSConv convolution decomposes the traditional convolution kernel into two components: a variable quantization kernel and a distribution offset, applies kernel-based and channel-based distribution offsets to maintain the same output as the original convolution, and achieves lower memory usage by storing only integer values, thereby achieving the purpose of improving model speed and reducing the number of parameters.

8. The method for intelligent track pavement damage segmentation based on image instance segmentation as claimed in claim 7, characterized in that: In step 2, during the quantization process of DSConv, the quantization function takes the number of bits to be quantized by the network as input, and uses 2's complement to save the integer value. For a number with a bit length of b, the following relationship exists: (1); In formula (1), Represents the value of each parameter in the tensor; First, the weights of each convolutional layer are scaled so that the original weights The maximum absolute value of matches the maximum value of the above quantization constraint; the new weight Stored in memory as integer values ​​for subsequent use in training and inference; by replacing ordinary convolution with DSConv, the memory saved for each tensor weight is: (2); By moving the VQK value through KDS and CDS, the weight of each block of the pre-trained network will be stretched or rounded to fit the interval in equation (1) and stored in VQK, where the optimal value of KDS is: (3); Its closed form is: , (4)。 9. The method for intelligent track pavement damage segmentation based on image instance segmentation as claimed in claim 8, characterized in that: The step 3 includes: the GUI interface includes four windows: image file selection, image display, recognition results and file saving.

Citation Information

Patent Citations

  • Pavement crack pixel level detection method based on instance segmentation algorithm

    CN112258529A

  • Ground surface crack instance segmentation method of multi-scale snakelike convolution constraint YOLOv8

    CN118674925A