Railway ballast track bed surface dirt quantification method and device based on machine vision
By combining the semantic segmentation model U-Net and the instance segmentation model SAM, the precise identification and quantification of the dirty surface of the ballast bed is achieved, and the comprehensive and data dependence problems of dirty evaluation in the prior art is solved, and the accurate assessment of the dirty rate of the dirty surface of the ballast bed is provided.
Patent Information
- Application Number
- CN202510355437.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
It is difficult for the prior art to effectively combine surface dirty data to conduct a comprehensive evaluation of the dirty condition of the ballast bed. The dirt evaluation effect of deep learning algorithms depends on the quality of the data set.
The image acquisition device is used to collect the track bed images, and the pre-trained semantic segmentation model U-Net and the instance segmentation big model SegmentAnything Model (SAM) are used to identify and segment dirty areas, and quantify them based on the area proportion of dirty areas.
It realizes the accurate identification and quantification of dirty diseases in large-scale road bed images, and provides an accurate assessment of the dirty rate of the road bed surface.
Smart Images

Figure CN120298338A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method and device for quantifying the surface dirt of a ballast bed of a railway based on machine vision. Background Art
[0002] In order to maintain the normal service of the ballast track, regular maintenance and repair are required, and the implementation of the maintenance plan depends on the accurate ballast pollution situation. In addition to traditional dirt assessment methods such as manual screening method, visual inspection method, and excavation method, non-destructive testing methods such as ground-penetrating radar, intelligent sensing ballast particles, infrared imaging, and resistivity are also used for on-site dirt identification. However, due to the diversity of dirt composition and distribution and the complexity inside the ballast bed, it has certain limitations to evaluate the dirt degree using any one of these technologies, and it is necessary to comprehensively evaluate the dirt situation of the ballast bed by combining surface dirt data, etc. As an important form of multi-source data, the apparent image can realize the rapid identification of the dirt area on the surface of the ballast bed of a railway, and provide support in terms of image data for the fusion evaluation system based on multi-source data. Deep learning algorithms show better effects compared with traditional image processing algorithms, but the dirt assessment effect based on deep learning algorithms is highly dependent on the quality of the dataset. Therefore, it is necessary to establish a method and device for quantifying the surface dirt of a ballast bed of a railway based on machine vision. Summary of the Invention
[0003] In order to solve the defects existing in the prior art, the present invention discloses a method for quantifying the surface dirt of a ballast bed of a railway based on machine vision, and its technical solution is as follows:
[0004] S1, using an image acquisition device to collect images of the ballast bed of a railway;
[0005] S2, inputting the obtained image into a pre-trained semantic segmentation model U-Net to identify and segment the dirt area in the ballast image, and obtaining a small-scale ballast dirt area;
[0006] S3, inputting the small-scale ballast dirt image into a new model obtained by fusing the semantic segmentation model U-Net and the instance segmentation large model SegmentAnything Model (SAM) to finely identify and segment the dirt area;
[0007] S4, according to the segmentation result of the ballast dirt image, using the area ratio of the dirt to obtain the surface dirt rate of the ballast bed.
[0008] Advantageous Effects
[0009] (1) Based on the semantic segmentation algorithm, the position of dirt diseases in a large-scale ballast bed image can be effectively obtained, which is beneficial to the subsequent quantification of the dirt rate;
[0010] (2)Based on the instance segmentation large model and the semantic segmentation model, the dirty area and the ballast area can be accurately segmented, and the surface dirt rate of the ballast bed is quantified by using the area ratio of the dirty area. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is the structural diagram of the semantic segmentation model U-Net in the present invention;
[0012] Figure 2 It is an example diagram of a large-scale ballast bed image of the present invention;
[0013] Figure 3 It is a schematic diagram of the ballast bed dirt annotation of the present invention;
[0014] Figure 4 It is a schematic diagram of the recognition effect of the dirty by the semantic segmentation model U-Net in the present invention;
[0015] Figure 5 It is the structural diagram of the instance segmentation large model SAM in the present invention;
[0016] Figure 6 It is a schematic diagram of the fusion recognition principle of U-Net and SAM in the present invention;
[0017] Figure 7 It is an example diagram of a small-scale dirty image dataset of the present invention;
[0018] Figure 8 It is a schematic diagram of the annotation of the small-scale dirty image of the present invention;
[0019] Figure 9 It is a schematic diagram of the segmentation result of the fusion model of SAM and U-Net in the present invention;
[0020] Figure 10 It is the linear regression of the surface dirt rate and the dirt mass ratio in the present invention;
[0021] Figure 11 It is a schematic flow diagram of quantifying the surface dirt degree of the ballast bed in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Next, the technical solutions in the embodiments of the present invention will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. These embodiments are provided to enable the present invention to be more thoroughly and completely conveyed to those skilled in the art.
[0023] The first aspect of the present invention provides a method for quantifying the surface dirt of a railway ballast bed based on machine vision, including the following steps:
[0024] S1, collecting ballast bed images of a ballast railway by using an image acquisition device;
[0025] Specifically, the effect of image data acquisition is as follows Figure 2 As shown, a large-scale ballast railway bed is photographed using an image acquisition device, and a multi-angle and multi-resolution shooting method is adopted.
[0026] S2. Input the obtained image into the pre-trained semantic segmentation model U-Net to identify and segment the dirty areas in the ballast image, and obtain a small-scale ballast dirty area;
[0027] Specifically, the network architecture of U-Net is as follows Figure 1 As shown. It mainly consists of two parts: a contracting path (reduction) and an expanding path (expansion). The contracting path is a typical convolutional neural network that extracts features through convolutional layers and pooling layers while reducing the data spatial dimension, while the expanding path restores the image spatial dimension through convolutional layers and upsampling layers, and combines low-dimensional and high-dimensional features at the same time to achieve pixel-level prediction. The contracting path and the expanding path are combined through "skip connections", as shown in Figure 1 As shown, the role of this connection is to splice the feature map with the corresponding feature map in the expanding path. It is very difficult to directly recover fine detail information from deep and high-level features. Therefore, skip connections can retain more details of the original input image, restore boundary information during the upsampling process, and thus improve the model performance.
[0028] When training the U-Net model, the energy function is calculated by combining the cross-entropy loss function with the per-pixel soft-max function on the final feature map. The soft-max function is defined as:
[0029]
[0030] where a k (x), a k' (x) represent the activation functions of feature channels k and k' at pixel position x ∈ Ω, respectively, and p k (x) is the approximate maximum function. That is, for k with the maximum activation a k (x), p k (x) ≈ 1; for all other k, p k (x) ≈ 0. The loss of cross-entropy at each position is expressed by the following formula:
[0031]
[0032] where Ω → {1,..., K} is the correct label of each pixel, w: Indicates the weight map obtained using pixels. The calculation formula for the weight map is:
[0033]
[0034] where w c : is the weight map that balances the class frequencies, d1: represents the distance to the nearest cell boundary; d2: represents the distance to the second nearest cell boundary. is the weight offset term, where w0 is regarded as a constant term representing the basic weight. According to experience, w0 is set to 10 and σ is about 5 pixels.
[0035] Furthermore, use the program to annotate the ballast dirt dataset, and only need to annotate the dirty and non-dirty areas. As Figure 3 shown, the annotated dirty areas do not consider the track foundation components such as sleepers and fasteners, and save the annotated images in the Microsoft COCO dataset format.
[0036] Furthermore, use the transforms module in the torchvision library to perform data augmentation on the original images and annotated images using operations such as translation, scaling, and color conversion, and increase the number of samples for model training to improve the generalization ability and accuracy of the model.
[0037] Specifically, the U-Net network can effectively identify and locate large-scale ballast dirt areas. The recognition effect of the model on the test set images is as Figure 4 shown.
[0038] S3, input the small-scale ballast dirt images into a new model obtained by fusing the semantic segmentation model U-Net and the instance segmentation large model SegmentAnything Model (SAM) to perform detailed recognition and segmentation of the dirty areas;
[0039] Specifically, SegmentAnything Model (SAM) is a new AI model released by MetaAI that can segment any object in an image. This model is based on a new architecture of Detection Transformer (DETR), which combines state-of-the-art object detection and segmentation techniques with the attention mechanism used in natural language. The SAM model mainly consists of 3 modules: an image encoding module, a prompt encoding module, and a mask decoding module. The model architecture is as Figure 5As shown in the figure. The structure of the SAM general model is not very complex. However, with an extremely large amount of training data, it is considered to have learned the general features of the pixel regions contained in an object. Therefore, even for a completely unfamiliar environment, it has good zero-shot segmentation performance, showing comparable or even slightly better results than well-trained supervised learning models. As the first promptable image segmentation base model, SAM was trained on the large-scale SA-1B dataset, which has an unprecedented number of images and annotations, enabling the model to have excellent zero-shot generalization ability.
[0040] In the SegmentAnything Model (SAM), a large model for deep learning instance segmentation, the image encoding module uses the vision transformer (VIT) architecture. Compared with convolutional networks, it is considered to be able to improve the upper limit of model training effect when the amount of data is large enough; the prompt encoding module uses convolution to encode prompts of the mask dense type, and position encoding is used for points and boxes; the mask decoding module efficiently combines the image embedding matrix output by the image encoding module and the encoded prompt information, and exports the mask corresponding to this prompt information. First, extract the geometric contour features of the particle contour. The pixel values of the segmented material contour are between 0 and 1. First, perform thresholding on the image, and then extract the area parameters respectively.
[0041] For the thresholding method described above, convert the exported single mask into a digital matrix and perform thresholding according to Equation (4);
[0042]
[0043] where pix(x, y) is the pixel value at the (x, y) coordinate position on the original image, and pix'(x, y) is the pixel at the (x, y) coordinate position on the processed image;
[0044] The calculation formula for the particle contour area is shown in Equation (5):
[0045] A = α 2 ·∑n x,y , x = 1, 2, 3, ···, w; y = 1, 2, 3, ···, h (5)
[0046] where A is the calculated area of the target contour, w is the image width, h is the image height, and n x,y represents the mapping value of the pixel. When the pixel value is 0, n x,y is 1, and α is the scale conversion factor, representing the actual length represented by a single pixel (unit: millimeters per pixel);
[0047] Input the construction waste image and a 32×32 grid of prompt points (the size of the prompt point grid can be adjusted according to the image size and the number of particles) into the instance segmentation large model, and a material mask with precise contour boundaries can be obtained.
[0048] Specifically, the fine image segmentation mask of SAM can be combined with the rich semantic annotations provided by the U-Net model and the two are fused. The new fused model is divided into three parts: a semantic branch, a mask branch, and a semantic voting module. The fused model is as shown in Figure 6 the figure. SAM acts as the mask branch, providing a mask with clear boundaries; the semantic branch does not need to provide very detailed boundaries, only needs to classify each region as accurately as possible. U-Net acts as the semantic branch, which provides a category for each pixel in the image, and the categories can be customized according to the network structure and requirements; the semantic voting module crops out the corresponding pixel categories according to the mask positions. Count the number of occurrences of each pixel category, and regard the category with the most occurrences as the classification result of the mask.
[0049] Furthermore, use a Python program to perform a closing operation on the identified dirty areas and crop the minimum bounding rectangle of the dirty areas.
[0050] Furthermore, small-scale dirty images taken on-site are as shown in Figure 7 (a). At the same time, in order to obtain more dirty images with different degrees of dirt and verify the algorithm, an indoor experiment is carried out. The acquisition of small-scale indoor dirty images is as shown in Figure 7 (b). Under the condition of consistent lighting conditions, use a mobile phone camera to take ballast dirty images with dirt contents of 5%, 10%, 15%, 20%, 25%, and 30% respectively. Use a program to annotate the images, only need to annotate the dirty areas and non-dirty areas. The image annotation is as shown in Figure 8 the figure. The rest can be regarded as the background area, and the data is saved in the Microsoft COCO dataset format.
[0051] Specifically, after training the fused model with a small-scale dataset, the segmentation effect can be obtained as shown in Figure 10 the figure.
[0052] S4. According to the segmentation results of the ballast dirty images, use the proportion of the dirty area to obtain the surface dirt rate of the ballast bed.
[0053] According to the segmentation results of the SAM and U-Net fused model, the dirt degree of the ballast image can be quantified by the proportion of the dirty area. The calculation of the surface dirt rate of the dirty area is as shown in Equation (6):
[0054]
[0055] Among them, FI_s represents the surface dirtiness rate, N f represents the number of dirty pixels, N a represents the total number of pixels in the image.
[0056] The surface dirtiness rate of the small - scale ballast image in the indoor test is calculated using Equation (6), and the surface dirtiness rates corresponding to the same dirt quality are averaged to obtain a one - to - one correspondence between the dirt quality ratio and the surface dirtiness rate.
[0057] Furthermore, the dirt quality ratio and the surface dirtiness rate are fitted, and the data with dirt quality of 10% and 20% are used for verification, as Figure 10 shown. It is not difficult to find that the two show a linear correlation, as shown in Equation (7).
[0058] FI_s = aω m +b (7)
[0059] Note: FI_s represents the surface dirtiness rate, ω m represents the proportion of dirt quality, and a and b are the fitting constants of the linear regression model.
[0060] Furthermore, the small - scale dirty image obtained by cropping the ballast bed image is input into the new fused model, and the surface dirtiness rate of this area can be obtained and mapped back to the original image.
[0061] The second aspect of the present invention provides a device for quantifying the surface dirtiness of railway ballast beds based on machine vision, including:
[0062] An acquisition module, used to vertically photograph the area within a large - scale ballast bed using a camera, including each component of the ballasted railway, to obtain a large - scale ballast bed image;
[0063] A positioning module, used to position the ballast bed image using a pre - trained semantic segmentation U - Net model to obtain a small - scale dirty area, and then crop out the minimum bounding rectangle of the dirty area;
[0064] A segmentation module, taking the small - scale ballast dirty image as input and using a new model fused with SAM and U - Net to perform detailed segmentation on it;
[0065] A quantification module, calculating the surface dirtiness rate of the area according to the detailed segmentation results of the dirty area and the non - dirty module, and mapping it back to the original image.
[0066] In summary, the present invention provides a method and device for quantifying the surface dirt of a railway ballast bed based on machine vision, including: identifying and locating the dirt disease areas in a large-scale track image based on the semantic segmentation model U-Net, and cropping the areas according to the minimum circumscribed rectangle of the dirt areas to obtain small-scale dirt images of different disease areas as the input for the next-stage dirt degree quantification algorithm. Then, a new model that fuses SAM and U-Net is used to finely segment the input small-scale dirt images, and the dirt degree is evaluated by the area ratio of the dirt area to the total area, and the dirt rate is mapped back to the original image to achieve the quantification of the surface dirt degree of the ballast bed. The process is as Figure 11 shown. This method can achieve accurate identification and segmentation of the ballast and dirt areas, providing a theoretical basis for the intelligent detection and quantitative identification of dirt diseases in ballast beds.
[0067] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision, characterized by: S1. Use an image acquisition device to collect images of the ballast bed of a ballast railway; S2. Input the acquired image into a pre-trained semantic segmentation U-Net model to identify and segment the dirty areas in the ballast image, and obtain a small-scale ballast dirty area; S3. Input the small-scale ballast dirty image into a new model obtained by fusing the semantic segmentation model U-Net and the instance segmentation large model SAM to conduct a detailed identification and segmentation of the dirty area; S4. According to the segmentation result of the ballast dirty image, use the proportion of the dirty area to obtain the surface dirt rate of the ballast bed.
2. The method for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision according to claim 1, characterized in that: Use a camera to vertically shoot the range of the ballast bed of a ballast railway. The captured image or video contains various railway components, and the dirty areas and non-dirty areas are marked in the captured image or video.
3. The method for quantifying the dirt on the surface of the ballast bed of a railway based on machine vision according to claim 1, wherein: Use computer vision algorithms to perform data augmentation on the training set, and conduct data training on the training set through transfer learning to obtain a semantic segmentation U-Net model. The training set is a large collection of ballast railway bed images captured in advance.
4. The method for quantifying the dirt on the surface of the ballast bed of a railway based on machine vision according to claim 2, wherein: Use the software annotation tool Labelme to annotate the ballast bed image training set.
5. The method for quantifying the dirt on the surface of the ballast bed of a railway based on machine vision according to claim 3, characterized in that: The network architecture of the semantic segmentation model U-Net includes a contracting path and an expanding path; the contracting path is a typical convolutional neural network, which conducts feature extraction through convolutional layers and pooling layers while reducing the data spatial dimension; the expanding path conducts image spatial dimension recovery through convolutional layers and upsampling layers, and combines low-dimensional and high-dimensional features at the same time to achieve pixel-level prediction.
6. The method for quantifying the dirt on the surface of the ballast bed of a railway based on machine vision according to claim 5, wherein: The contracting path and the expanding path splice the feature maps with the corresponding feature maps in the expanding path.
7. The method for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision according to claim 6, characterized in that: training When using the semantic segmentation model U-Net, the energy function is calculated by combining the cross-entropy loss function with the per-pixel soft-max function on the final feature map: the soft-max function is defined as: where a k (x), a k' (x) represent the activation functions of the feature channels k and k' at the pixel position x ∈ Ω, respectively; K is the number of analogies, p k (x) is the approximate maximum function; that is, for k with the maximum activation a k (x), p k (x) ≈ 1; for all other k, p k (x) ≈ 0; the loss p l(x) (x) at each position is expressed by the following formula: where \(l:\Omega\rightarrow\{1,\ldots,K\}\) is the correct label for each pixel, denotes the weight map obtained using the pixels.
8. The method for quantifying the dirt on the surface of the ballast bed of a railway based on machine vision according to claim 7, characterized in that: The calculation formula of the weight map is: wherein is a weight map for balancing class frequencies, expressed as the distance to the nearest cell boundary; expressed as the distance to the second nearest cell boundary. is the weight offset term, where w0 is regarded as a constant term representing the basic weight. According to experience, w0 is set to 10 and σ is about 5 pixels.
9. The method for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision according to claim 1, characterized in that: According to the segmentation result of the semantic segmentation model, the minimum bounding rectangle of the dirty area is used to crop the area to obtain small-scale dirty images of different disease areas and input them into the new model obtained by fusing the instance segmentation large model SAM and U-Net in the next stage.
10. The method for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision according to claim 9, characterized in that: For the deep learning instance segmentation large model SAM used, for the prompts of the mask dense type, convolution is used for encoding, and points and boxes use position encoding; The mask decoding module efficiently combines the image embedding matrix output by the image encoding module and the encoded prompt information, and exports the mask corresponding to this prompt information.
11. The method for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision according to claim 10, characterized in that: Use the SAM model to first extract the geometric contour features of the particle contour. The pixel values of the segmented material contour are between 0 and 1. First, perform thresholding on the image, and then extract the area parameters respectively.
12. The method for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision according to claim 11, characterized in that: The thresholding method is to convert the exported single mask into a digital matrix and perform thresholding according to Equation (4); where pix(x, y) is the pixel value at the coordinate position (x, y) on the original image, and pix'(x, y) is the pixel at the coordinate position (x, y) on the processed image; The calculation formula for the particle contour area is as shown in Equation (5): A = α 2 ·∑n x,y , x = 1, 2, 3, …, w; y = 1, 2, 3, …, h (5) Where A is the calculated area of the target contour, w is the image width, h is the image height, and n x,y represents the mapping value of the pixel. When the pixel value is 0, n x,y is 1, α is the scale conversion factor, representing the actual length represented by a single pixel, unit: millimeters per pixel.
13. The method for quantifying the dirt on the surface of the ballast bed of a railway based on machine vision according to claim 9, characterized in that: The new model obtained by fusing the instance segmentation large model SAM and U-Net is divided into three parts: a semantic branch, a mask branch, and a semantic voting module.
14. The method for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision according to claim 1, characterized in that: According to the segmentation results of the SAM and U-Net fusion model, the degree of soiling of the ballast image can be quantified by the proportion of the area of the soiled area; the calculation of the surface soiling rate of the soiled area is shown in Equation (6): Among them, FI_s represents the surface dirt rate, N f represents the number of dirty pixels, N a represents the total number of pixels in the image.
15. The ballast bed surface dirt quantification device for railways based on machine vision is characterized in that, Including: An acquisition module for vertically photographing an area within a large range of the ballast bed using a camera, including components of each ballasted railway, to obtain a large-range ballast bed image; A positioning module for using a pre-trained semantic segmentation U-Net model to position the ballast bed image to obtain a small-range soiled area, and then cropping out the minimum bounding rectangle of the soiled area; A segmentation module that takes the small-range ballast soiled image as input and uses the new model fused with SAM and U-Net to perform detailed segmentation on it; A quantification module that calculates the surface soiling rate of the area based on the detailed segmentation results of the soiled area and the non-soiled module, and maps it back to the original image.
16. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, wherein the program, when running, controls the device where the non-volatile storage medium is located to execute the method for efficient classification and component ratio estimation of construction waste according to any one of claims 1 to 14.
17. The electronic device for quantifying the dirt on the surface of a ballast bed of a railway based on machine vision is characterized in that, It includes a processor and a memory; computer-readable instructions are stored in the memory, and the processor is used to run the computer-readable instructions, wherein the computer-readable instructions, when running, execute the method for quantifying the surface soiling of a ballasted railway bed based on machine vision according to any one of claims 1 to 14.
Citation Information
Patent Citations
Rail foreign matter detection method and system under space-based visual angle based on weak supervised learning
CN111582084A
Remote sensing image-based cultivated land extraction method, device, equipment and medium
CN116994140A
Natural element extraction method based on high-resolution remote sensing image
CN117392550A
Small-sample dam crack segmentation method and system based on general segmentation large model
CN118736226A
Semantic segmentation method and device based on SAM model, equipment and storage medium
CN119206207A
Cited By
Multi-dyeing livestock and poultry muscle fiber quantification method and system based on double-model fusion
CN121304668A
Quantitative method and system for multiple staining of livestock and poultry muscle fibers based on double model fusion
CN121304668B
Leakage magnetic field scanning steel bar corrosion image recognition system and method thereof
CN121391877A
Magnetic flux leakage field scanning steel bar corrosion image recognition system and method thereof
CN121391877B
Method for detecting sand content of sandstorm ballast bed based on self-adaptive semantic segmentation
CN121661394A