Railway vehicle crack detection method and system based on machine vision

Through the machine vision-based crack detection method of rail vehicle, the limitations of the detection object structure and the difficulty of identifying fine cracks in the prior art are solved, and timely identification, precise positioning and classification of cracks are realized, and detection accuracy and accuracy are improved.

CN120198404APending Publication Date: 2025-06-24HUNAN FIRST NORMAL UNIV

Patent Information

Application Number
CN202510339018.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing rail vehicle crack detection technology has limitations on the structure of the detection object, making it difficult to identify fine cracks and cannot be accurately classified.

Method used

Using machine vision-based detection method, motion blur compensation and multi-frame image fusion are performed by acquiring continuous images of the rail vehicle surface in a moving state, combining adaptive contrast enhancement and improved histogram equalization, crack candidate regions are located, and feature extraction and classification are used using improved ResNet deep learning network and attention mechanism.

Benefits of technology

It realizes timely identification and precise positioning of fine cracks, can improve detection accuracy in high-speed motion scenarios, and can accurately classify crack types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198404A_ABST
    Figure CN120198404A_ABST
Patent Text Reader

Abstract

The invention discloses a railway vehicle crack detection method and system based on machine vision, and the method comprises the steps: S1, obtaining continuous images of the surface of a railway vehicle in a motion state, carrying out the motion blur compensation, carrying out the fusion of the compensated images, and obtaining a vehicle surface fusion image; s2, performing adaptive contrast enhancement and improved histogram equalization processing on the vehicle surface fusion image to obtain a processed image; s3, performing crack candidate region positioning on the processed image based on adaptive threshold segmentation and connected domain analysis; s4, utilizing an improved ResNet deep learning network to perform feature extraction on the positioned crack candidate region, and constructing a crack feature vector set; and S5, classifying the crack feature vectors by adopting a multi-scale fusion algorithm based on an attention mechanism, and outputting positions and category results of the cracks. According to the method, the histogram equalization and the learning network are improved to identify the fine cracks and realize accurate positioning and classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crack detection, and in particular, to a method and system for detecting cracks in rail vehicles based on machine vision. Background Art

[0002] With the rapid development of high-speed railways, the real-time detection and classification technology of cracks on the surface of rail vehicles is of great significance for ensuring train operation safety. During the high-speed operation of rail vehicles, different types of cracks are likely to occur on their surfaces due to factors such as fatigue and impact. If these cracks cannot be detected and handled in a timely manner, potential safety hazards will be generated for train operation.

[0003] Currently, the crack detection of rail vehicles mainly includes manual visual inspection method and machine vision inspection method. The manual visual inspection method uses manual patrol to visually inspect the vehicle surface; however, this method has low efficiency, high labor cost, unstable detection level and is easily affected by the subjective factors of inspectors, making it difficult to meet the all-weather and high-frequency detection requirements. The traditional machine vision inspection method uses a fixed camera to capture images of the vehicle surface and perform image processing; however, this method is prone to image blurring problems in high-speed motion scenarios, affecting the detection accuracy, and the camera shooting is sensitive to lighting conditions, and its performance is unstable in complex train operation environments.

[0004] The existing patent CN202411848175.8 discloses a crack detection system for device based on millimeter-wave imaging radar. This system uses point cloud data modeling and calculates deviation factors based on the symmetry principle for crack detection. Although this system has certain environmental adaptability, it still has the following defects: First, it requires that the detection object must have symmetric structural features, which greatly limits the application scenarios and scope; second, the spatial resolution of point cloud data is relatively low, making it difficult to identify tiny cracks and unable to achieve timely monitoring of cracks; third, it can only detect the crack position and cannot accurately classify the crack types. Summary of the Invention

[0005] Based on the above problems, the purpose of the present invention is to provide a new method for detecting cracks in rail vehicles, which can get rid of the limitation of the structure of the detection object, realize the timely identification of fine cracks, and the accurate positioning and classification of cracks.

[0006] To achieve the above purpose, the present invention discloses a method for detecting cracks in rail vehicles based on machine vision, including the following steps:

[0007] S1: Obtain continuous images of the surface of a rail vehicle in a moving state and perform motion blur compensation, and fuse the compensated images to obtain a fused image of the vehicle surface;

[0008] S2: Perform adaptive contrast enhancement and improved histogram equalization on the fused image of the vehicle surface to obtain the processed image; the improved histogram equalization includes introducing contrast limitation and histogram redistribution during the calculation of the cumulative distribution function.

[0009] S3: Locate the crack candidate regions in the processed image based on adaptive threshold segmentation and connected component analysis.

[0010] S4: Use the improved ResNet deep learning network to extract features from the located crack candidate regions and construct a crack feature vector set; the improved ResNet deep learning network includes introducing channel and spatial attention modules to adaptively focus on the key regions and important feature channels of the cracks.

[0011] S5: Classify the crack feature vectors using a multi-scale fusion algorithm based on the attention mechanism and output the position and category results of the cracks.

[0012] Preferably, the step S1 includes the following steps:

[0013] S11: Use a high-speed camera to collect a continuous multi-frame image of the surface of the rail vehicle to obtain an image sequence I = {I1, …, I t , …, I N}, where I t is the t-th frame image, t is the frame number, N is the total number of frames in the image sequence, and the value range is t ∈ [1, N];

[0014] S12: Perform motion estimation on adjacent frames in the image sequence and calculate the motion vector field M t :

[0015]

[0016] where M t (x, y) is the motion vector of the t-th frame image at the pixel coordinate (x, y), dx and dy are the displacements in the horizontal and vertical directions respectively, w is the search window size, i and j are the offsets in the horizontal and vertical directions within the search window, I t (x + i, y + j) is the pixel value of the t-th frame image at the pixel coordinate (x + i, y + j), I t-1 (x + dx + i, y + dy + j) is the pixel value of the (t - 1)-th frame image at the pixel coordinate (x + dx + i, y + dy + j), and ‖·‖2 is the L2 norm;

[0017] S13: Perform multi-frame image fusion based on the motion vector field to obtain the compensated fused image I fusion .

[0018] Preferably, the step S13 includes the following steps:

[0019] S131: Perform a spatial transformation on each frame of image according to the motion vector field to obtain a compensated single-frame image:

[0020]

[0021] Wherein, is the pixel value at the pixel coordinate (x, y) of the compensated image of the t-th frame, I is the pixel value of the original image of the t-th frame at the displaced coordinate, t and and are respectively the horizontal component and the vertical component of the motion vector field M t at the pixel coordinate (x, y), is the floor operation;

[0022] S132: Perform multi-frame image fusion on the compensated single-frame image to obtain a vehicle surface fusion image:

[0023]

[0024] Wherein, I fusion (x, y) is the pixel value of the vehicle surface fusion image I fusion at the pixel coordinate (x, y), W t (x, y) is the compensated image of the t-th frame at the pixel coordinate (x, y), the weight coefficient, is the compensated image of the t-th frame at the pixel coordinate (x, y), the gradient value, σ is the weight decay speed control parameter, and e is the natural constant.

[0025] Preferably, the step S2 includes the following steps:

[0026] S21: Perform adaptive local contrast enhancement on the vehicle surface fusion image:

[0027] I contrast (x, y) = μ local (x, y) + β(x, y) × (I fusion (x, y) - μ local (x, y))

[0028]

[0029] Wherein, I contrast (x, y) is the enhanced image I contrastThe pixel value at pixel coordinates (x, y), μ local (x, y) is the vehicle surface fusion image I fusion The pixel mean value within a 5×5 local region centered at pixel coordinates (x, y), β(x, y) is the adaptive gain coefficient, γ is the global gain coefficient, σ global is the vehicle surface fusion image I fusion The global pixel standard deviation of, σ local (x, y) is the vehicle surface fusion image I fusion The pixel standard deviation within a 5×5 local region centered at pixel coordinates (x, y);

[0030] S22: Perform improved histogram equalization processing on the enhanced image I contrast

[0031] Preferably, the step S22 includes the following steps:

[0032] S221: Calculate the cumulative distribution function:

[0033]

[0034] where level is the gray value, level ∈ [0, L - 1], L is the number of gray levels, and h(k, x, y) is the number of pixels with gray value k within a 5×5 local region centered at pixel coordinates (x, y);

[0035] S222: Contrast limitation and histogram redistribution:

[0036]

[0037] where, is the limited histogram value, α is the contrast limitation coefficient, and E(x, y) is the number of pixels exceeding the limit;

[0038] Average the pixel excess part to all gray levels to obtain the redistributed histogram:

[0039]

[0040] S223: Obtain the processed image:

[0041] Update the cumulative distribution function:

[0042]

[0043] Use the updated cumulative distribution function to calculate and obtain the processed image:

[0044] I process (x, y) = (L - 1) × CDF​new (I contrast (x, y), x, y)

[0045] Among them, I process (x, y) is the pixel value of the processed image I process at the pixel coordinates (x, y).

[0046] Preferably, the step S3 includes the following steps:

[0047] S31: Calculate the local adaptive threshold:

[0048]

[0049] Among them, T(x, y) is the local adaptive threshold at the pixel coordinates (x, y), is the average pixel value within the 7×7 local area centered on the pixel coordinates (x, y) of the processed image I process , and C is a fixed offset;

[0050] S32: Binarize the processed image:

[0051]

[0052] Among them, I binary (x, y) is the pixel value of the binary image I binary at the pixel coordinates (x, y);

[0053] S33: Perform morphological processing on the binary image:

[0054] I morph = Close(Open(I binary , K1), K2)

[0055] Among them, I morph is the result of morphological processing, Open(·) and Close(·) are the opening operation and closing operation respectively, and K1 and K2 are the morphological operator kernels;

[0056] S34: Connected component analysis and geometric feature screening:

[0057]

[0058] Among them, {R p′} is the set of connected components, N′ is the number of connected components, ConnectedComponents(·) is the connected component analysis function, λ p′ and ω p′ are the length and width of the p′-th connected component respectively, and A pis the area of the p'-th connected region, τ1 is the aspect ratio threshold, τ2 is the area threshold, and R crack is the set of crack connected regions after screening;

[0059] S35: Calculate the minimum bounding rectangle as the crack candidate region:

[0060]

[0061] where, and are the abscissa and ordinate of the center of the candidate region respectively, and are the width and height of the candidate region respectively, MinBoundingRect(·) is the minimum bounding rectangle calculation function, and R p is the p-th crack connected region.

[0062] Preferably, the step S4 includes the following steps:

[0063] S41: Normalize the size of each crack candidate region:

[0064]

[0065] where, is the image of the p-th normalized crack candidate region, Crop(I process ,B p ) is to crop the region bounded by B process from the processed image I p , is the bilinear interpolation operation to adjust the input image to the standard size, and are the normalized width and height respectively;

[0066] S42: Construct an improved ResNet deep learning network:

[0067]

[0068] where, F1, F2, and F3 are the feature maps obtained after their respective operations, Conv(·,θ1) is the convolution operation function based on the parameter θ1, ResBlock(·,θ2) is the residual block operation function based on the parameter θ2, and CBAM(·,θ3) is the channel and spatial attention module based on the parameter θ3, is the number of channels of the feature map F1, is the number of channels of the feature maps F2 and F3, is the height of the feature map F1, is the height of the feature maps F2 and F3, is the width of the feature map F1, is the width of the feature maps F2 and F3, is the real number field;

[0069] S43: Extract the feature vectors and construct a crack feature vector set:

[0070]

[0071] V = {v p |R p ∈R crack}

[0072] where GAP(·) is the global average pooling operation function, v p is the feature vector of the p-th crack candidate region, and V is the crack feature vector set composed of the feature vectors of all crack candidate regions.

[0073] Preferably, the step S5 includes the following steps:

[0074] S51: Construct a multi-scale feature pyramid:

[0075] Q i′ = DownSample(V, ε i′ ), i′ ∈ {1, 2, 3}

[0076]

[0077] where DownSample(·, ε i′ ) is a function for downsampling the feature vectors based on the i'-th downsampling scale factor ε i′ , UpSample(·, δ i′ ) is a function for upsampling based on the i'-th upsampling scale factor δ i′ , Q i′ is the feature of the i'-th layer of the feature pyramid, is the feature after upsampling Q i′ to a unified scale;

[0078] S52: Calculate the attention weights:

[0079]

[0080] where SoftAttention(·, ψ) is a function for calculating the soft attention weights based on the parameter ψ, and φ i′ is the soft attention weight of the i'-th layer of the feature;

[0081] S53: Feature fusion and classification:

[0082]

[0083] Among them, Z is the fused feature vector, is a function for performing a fully connected layer calculation based on the parameter , Softmax(·) is a normalization function, Prob(cls|Z) is the class probability distribution under the condition of the given feature Z, cls ∈ {0, 1, …, η} is the crack class, and η is the total number of crack classes;

[0084] S54: Output of crack detection result:

[0085]

[0086] G = {(B p , cls * )}

[0087] Among them, cls * is the predicted crack class, G is the set of final crack detection results, including the crack position and the crack class.

[0088] The present invention also discloses an on-rail vehicle crack detection system based on machine vision, including:

[0089] Image preprocessing module: Obtain continuous images of the surface of an on-rail vehicle in a moving state and perform motion blur compensation, and fuse the compensated images to obtain a fused image of the vehicle surface;

[0090] Image enhancement module: Perform adaptive contrast enhancement and improved histogram equalization processing on the fused image of the vehicle surface to obtain a processed image;

[0091] Crack candidate region positioning module: Locate crack candidate regions in the processed image based on adaptive threshold segmentation and connected component analysis;

[0092] Feature extraction module: Use an improved ResNet deep learning network to extract features from the located crack candidate regions and construct a set of crack feature vectors;

[0093] Classification and output module: Classify the crack feature vectors using a multi-scale fusion algorithm based on an attention mechanism and output the crack position and class results.

[0094] Compared with the prior art, the present invention has at least the following beneficial effects:

[0095] The present invention first uses a spatial transformation method based on the motion vector field to perform motion compensation on adjacent frame images, effectively eliminating motion blur. Then, through the collaborative fusion of multiple frame images and further introducing an adaptive weight coefficient based on image gradients, the fusion process can highlight the detailed information of the images, significantly improving the image quality and signal-to-noise ratio, and laying a good foundation for subsequent crack detection.

[0096] The present invention processes the fused image by introducing an adaptive local contrast enhancement algorithm, which can dynamically adjust the gain coefficient according to the statistical characteristics of the local area of the image, effectively solving the problem of uneven illumination in different regions. The present invention improves the histogram equalization technology by introducing a contrast limit and redistribution mechanism in the calculation process of the cumulative distribution function, effectively suppressing noise amplification while enhancing the image contrast and maintaining the naturalness of the image.

[0097] The present invention further enhances the feature extraction ability based on the improved ResNet network structure and CBAM attention module. By constructing a multi-scale feature pyramid, the present invention can simultaneously capture the local details and global structural features of cracks, and adaptively adjust the weights of features at different scales by introducing a soft attention mechanism, achieving the optimal fusion of features. Description of the Drawings

[0098] Figure 1 is a flowchart of a method for detecting cracks in rail vehicles based on machine vision according to Embodiment 1 of the present invention;

[0099] Figure 2 is an example diagram for locating crack candidate regions in Embodiment 1 of the present invention, where Fig. (a) is a local image including cracks, and Fig. (b) is an image of the located crack candidate region. Detailed Embodiments

[0100] The present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited in any way. Any transformation or replacement made based on the teachings of the present invention falls within the protection scope of the present invention.

[0101] Embodiment 1:

[0102] As Figure 1 shown, a method for detecting cracks in rail vehicles based on machine vision includes the following steps:

[0103] S1: Obtain continuous images of the surface of a rail vehicle in a moving state and perform motion blur compensation, and fuse the compensated images to obtain a fused image of the vehicle surface; specifically:

[0104] S11: Use a high-speed camera to collect continuous multiple frames of images of the surface of the rail vehicle to obtain an image sequence I = {I1,..., I t, …, I N}, where I t is the t-th frame image, t is the frame sequence number, N is the total number of frames in the image sequence, and the value range is t ∈ [1, N];

[0105] S12: Perform motion estimation on adjacent frames in the image sequence, and calculate the motion vector field M t :

[0106]

[0107] Among them, M t (x, y) is the motion vector of the t-th frame image at the pixel coordinates (x, y), dx and dy are the displacements in the horizontal and vertical directions respectively, w is the size of the search window, which is 7 in this embodiment, i and j are the offsets in the horizontal and vertical directions within the search window, and I t (x + i, y + j) is the pixel value of the t-th frame image at the pixel coordinates (x + i, y + j), and I t-1 (x + dx + i, y + dy + j) is the pixel value of the (t - 1)-th frame image at the pixel coordinates (x + dx + i, y + dy + j), and ‖·‖2 is the L2 norm;

[0108] S13: Perform multi-frame image fusion based on the motion vector field to obtain the compensated vehicle surface fusion image I fusion ; specifically:

[0109] S131: Perform spatial transformation on each frame image according to the motion vector field to obtain the compensated single-frame image:

[0110]

[0111] Among them, is the pixel value of the compensated t-th frame image at the pixel coordinates (x, y), and I t is the pixel value of the original t-th frame image at the displaced coordinates, and are the horizontal and vertical components of the motion vector field M t at the pixel coordinates (x, y) respectively, is the floor operation;

[0112] S132: Perform multi-frame image fusion on the compensated single-frame images to obtain the vehicle surface fusion image:

[0113]

[0114] Among them, I fusion(x, y) is the vehicle surface fusion image I fusion The pixel value at pixel coordinates (x, y), W t (x, y) is the image after compensation for the t-th frame The weight coefficient at pixel coordinates (x, y), is the image after compensation for the t-th frame The gradient value at pixel coordinates (x, y), σ is the weight decay rate control parameter, which is 25 in this embodiment, and e is the natural constant.

[0115] In this step, through the multi-frame image processing method, the image quality problem in the high-speed motion scene is effectively solved. First, a high-speed camera is used to obtain a series of consecutive frames of images, providing basic data for subsequent processing; second, through an accurate motion estimation algorithm, a pixel-level correspondence is established between adjacent image frames, realizing the accurate calculation of the motion vector field. This motion estimation method not only improves the accuracy of motion estimation by minimizing pixel differences within a local search window but also has strong anti-noise ability.

[0116] S2: Perform adaptive contrast enhancement and improved histogram equalization processing on the vehicle surface fusion image to obtain the processed image; specifically:

[0117] S21: Perform adaptive local contrast enhancement on the vehicle surface fusion image:

[0118] I contrast (x, y) = μ local (x, y) + β(x, y) × (I fusion (x, y) - μ local (x, y))

[0119]

[0120] Where, I contrast (x, y) is the enhanced image I contrast The pixel value at pixel coordinates (x, y), μ local (x, y) is the vehicle surface fusion image I fusion The pixel mean within a 5×5 local area centered on pixel coordinates (x, y), β(x, y) is the adaptive gain coefficient, γ is the global gain coefficient, which is 1.5 in this embodiment, σ global is the vehicle surface fusion image I fusion The global pixel standard deviation of, local (x, y) is the vehicle surface fusion image I fusion The pixel standard deviation within a 5×5 local area centered on pixel coordinates (x, y);

[0121] S22: For the enhanced image Icontrast Perform improved histogram equalization processing; specifically:

[0122] S221: Calculate the cumulative distribution function:

[0123]

[0124] where level is the gray value, level ∈ [0, L - 1], L is the number of gray levels, and h(k, x, y) is the number of pixels with gray value k in the 5×5 local region centered at the pixel coordinates (x, y);

[0125] S222: Contrast limitation and histogram redistribution:

[0126]

[0127] where is the histogram value after limitation, α is the contrast limitation coefficient, which is 40 in this embodiment, and E(x, y) is the number of pixels exceeding the limitation;

[0128] Distribute the excess part of the pixels evenly to all gray levels to obtain the redistributed histogram:

[0129]

[0130] S223: Obtain the processed image:

[0131] Update the cumulative distribution function:

[0132]

[0133] Calculate and obtain the processed image using the updated cumulative distribution function:

[0134] I process (x, y) = (L - 1) × CDF new (I contrast (x, y), x, y)

[0135] where I process (x, y) is the pixel value of the processed image I process at the pixel coordinates (x, y).

[0136] In this step, in the contrast enhancement stage, an adaptive gain control mechanism based on local statistical characteristics is adopted. By dynamically adjusting the gain coefficient, a larger gain is obtained for the dark regions of the image while a smaller gain is obtained for the bright regions, thereby improving the overall visibility of the image while avoiding over-enhancement and noise amplification. Especially when processing vehicle surface images with uneven illumination, this method can adaptively adjust the enhancement degree according to the standard deviation of the local region, ensuring the local optimality of the enhancement effect; in the histogram equalization stage, a contrast limitation and histogram redistribution mechanism are introduced. By setting reasonable limitation thresholds and redistribution strategies, the over-enhancement and pseudo-contour phenomena easily caused by traditional histogram equalization are effectively suppressed.

[0137] S3: Locate the crack candidate regions in the processed image based on adaptive threshold segmentation and connected component analysis; specifically:

[0138] S31: Calculate the local adaptive threshold:

[0139]

[0140] Among them, T(x, y) is the local adaptive threshold at the pixel coordinates (x, y), is the mean value of the pixels within the 7×7 local region centered at the pixel coordinates (x, y) of the processed image I process C is a fixed offset, which is -15 in this embodiment;

[0141] S32: Perform binarization processing on the processed image:

[0142]

[0143] Among them, I binary (x, y) is the pixel value of the binary image I binary at the pixel coordinates (x, y);

[0144] S33: Perform morphological processing on the binary image:

[0145] I morph = Close(Open(I binary , K1), K2)

[0146] Among them, I morph is the result of morphological processing, Open(·) and Close(·) are the opening operation and closing operation respectively, and K1 and K2 are the morphological operator kernels. In this embodiment:

[0147]

[0148] S34: Connected component analysis and geometric feature screening:

[0149]

[0150] Among them, {R p′} is the set of connected regions, N′ is the number of connected regions, ConnectedComponents(·) is the connected region analysis function, λ p′ and ω p′ are the length and width of the p′-th connected region respectively, A p is the area of the p′-th connected region, τ1 is the aspect ratio threshold, which is 3.0 in this embodiment, τ2 is the area threshold, which is 50 in this embodiment, and R crack is the set of screened crack connected regions;

[0151] S35: Calculate the minimum bounding rectangle as the crack candidate region:

[0152]

[0153] Among them, and are the abscissa and ordinate of the center of the candidate region respectively, and are the width and height of the candidate region respectively, MinBoundingRect(·) is the minimum bounding rectangle calculation function, and R p is the p-th crack connected region.

[0154] In this step, the local adaptive threshold segmentation method is adopted, which fully considers the brightness change characteristics of the local area of the image. Compared with the global threshold method, it has stronger environmental adaptability and can effectively cope with the influence brought by the uneven illumination on the vehicle surface and the complex background; in addition, on the basis of binaryzation processing, through morphological processing operations, including the combined application of opening operation and closing operation, the image noise and non-target small regions are effectively suppressed, while maintaining the integrity and continuity of the crack target.

[0155] Refer to Figure 2 the example given. Through this step, the region where the crack is located in the image with cracks is further reduced until the smallest crack candidate region is framed, which is beneficial to the accurate positioning of the crack position.

[0156] S4: Use the improved ResNet deep learning network to extract features from the located crack candidate regions and construct a set of crack feature vectors; specifically:

[0157] S41: Normalize the size of each crack candidate region:

[0158]

[0159] Among them, is the p-th normalized crack candidate region image, Crop(I process , B p ) is to crop the region bounded by B process from the processed image I p , is the bilinear interpolation operation to adjust the input image to the standard size, and are the normalized width and height respectively, both being 224 in this embodiment;

[0160] S42: Construct an improved ResNet deep learning network:

[0161]

[0162] Among them, F1, F2, and F3 are the feature maps obtained after their respective operations, Conv(·, θ1) is the convolution operation function based on the parameter θ1, ResBlock(·, θ2) is the residual block operation function based on the parameter θ2, and CBAM(·, θ3) is the channel and spatial attention module based on the parameter θ3, is the number of channels of the feature map F1, is the number of channels of the feature maps F2 and F3, is the height of the feature map F1, is the height of the feature maps F2 and F3, is the width of the feature map F1, is the width of the feature maps F2 and F3, is the real number field;

[0163] In this embodiment, the specific parameter configuration of the network structure is as follows:

[0164] Input layer: Use bilinear interpolation to adjust the crack candidate region to the standard size of 224×224;

[0165] First convolutional layer: Use 64 3×3 convolutional kernels with a stride of 2, and the size of the output feature map after convolution operation is 64×112×112;

[0166] Residual block: It contains 3 sub-modules, each sub-module consists of two 3×3 convolutional layers, and the channel is adjusted through 1×1 convolution, and finally 256 channels are output;

[0167] CBAM attention module: The channel attention branch uses a combination of adaptive average pooling and max pooling, and the spatial attention branch uses a 7×7 convolutional kernel;

[0168] The size of the feature maps output after the residual block operation and the attention module processing is 256×56×56;

[0169] S43: Extract the feature vectors and construct a crack feature vector set:

[0170]

[0171] V = {v p |R p ∈R crack}

[0172] where GAP(·) is the global average pooling operation function, v p is the feature vector of the p-th crack candidate region, and V is the crack feature vector set composed of the feature vectors of all crack candidate regions.

[0173] In this step, first, by performing size normalization on the crack candidate regions, the consistency of the input data is ensured; second, an improved ResNet network structure is designed. By introducing channel and spatial attention modules, it can adaptively focus on the key regions and important feature channels of the cracks, significantly improving the pertinence and discriminability of feature extraction; finally, global average pooling operation is used to generate feature vectors, which not only reduces the feature dimension but also maintains the spatial invariance of the features, making the extracted features have better generalization ability.

[0174] S5: Classify the crack feature vectors using a multi-scale fusion algorithm based on the attention mechanism, and output the position and category results of the cracks; specifically:

[0175] S51: Construct a multi-scale feature pyramid:

[0176] Q i′ = DownSample(V, ε i′ ), i′ ∈ {1, 2, 3}

[0177]

[0178] where DownSample(·, ε i′ ) is the function for downsampling the feature vectors based on the i′-th downsampling scale factor ε i′ . In this embodiment, ε1 = 0.5, ε2 = 0.25, ε3 = 0.125, and UpSample(·, δ i′ ) is the function for upsampling based on the i′-th upsampling scale factor δ i′ . In this embodiment, δ1 = 2, δ2 = 4, δ3 = 8, Q i′ is the feature of the i′-th layer of the feature pyramid, is the feature after upsampling Q i′ to a unified scale;

[0179] S52: Calculate attention weights:

[0180]

[0181] where SoftAttention(·, ψ) is a function for calculating soft attention weights based on parameter ψ, and φ i′ is the soft attention weight of the features at the i'-th layer;

[0182] In this embodiment, the configuration of the attention module is as follows:

[0183] The input feature dimension is 256; the attention intermediate layer dimension is 128; the attention calculation uses a two-layer perceptron structure; the activation function is ReLU; the Softmax temperature coefficient is set to 1.0;

[0184] S53: Feature fusion and classification:

[0185]

[0186] where Z is the fused feature vector, is a function for performing fully connected layer calculations based on parameter In this embodiment, the number of hidden units in the fully connected layer is 1024, Softmax(·) is a normalization function, Prob(cls|Z) is the class probability distribution given feature Z, cls ∈ {0, 1,..., η} is the crack class, and η is the total number of crack classes. In this embodiment, η = 4, including normal samples, transverse cracks, longitudinal cracks, and reticular cracks;

[0187] S54: Output of crack detection results:

[0188]

[0189] G = {(B p , cls * )}

[0190] where cls * is the predicted crack class, and G is the final crack detection result set, which contains the crack location and the crack class.

[0191] In this step, by constructing a multi-scale feature pyramid, the system systematically captures the performance of crack features at different scales. The design of the feature pyramid realizes the effective extraction of multi-scale features through downsampling and upsampling operations while maintaining the integrity of feature information. Secondly, the soft attention mechanism introduced in this step can adaptively assign weights to features at different scales, enabling the system to intelligently focus on the most discriminative feature levels and significantly improving the classification accuracy.

[0192] Embodiment 2:

[0193] An on-vehicle track crack detection system based on machine vision includes the following five modules:

[0194] Image preprocessing module: Obtain continuous images of the surface of an on-vehicle track in motion and perform motion blur compensation, and fuse the compensated images to obtain a fused image of the vehicle surface;

[0195] Image enhancement module: Perform adaptive contrast enhancement and improved histogram equalization processing on the fused image of the vehicle surface to obtain a processed image;

[0196] Crack candidate region localization module: Locate crack candidate regions in the processed image based on adaptive threshold segmentation and connected component analysis;

[0197] Feature extraction module: Use an improved ResNet deep learning network to extract features from the located crack candidate regions and construct a crack feature vector set;

[0198] Classification and output module: Classify crack feature vectors using a multi-scale fusion algorithm based on an attention mechanism and output the position and category results of the cracks.

[0199] The crack detection system provided in this embodiment is used to implement the crack detection method in the above-mentioned Embodiment 1. Among them, the functions realized by each functional module in the crack detection system correspond one by one to each process step in the crack detection method; therefore, it will not be elaborated here.

[0200] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments. And the term "including", "comprising" or any other variant thereof in this article is intended to cover a non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, device, article or method including the element.

[0201] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0202] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for detecting cracks in rail vehicles based on machine vision, characterized in that: The following steps are involved: S1: acquiring continuous images of the surface of a moving rail vehicle and performing motion blur compensation, fusing the compensated images to obtain a fused image of the vehicle surface; S2: performing adaptive contrast enhancement and improved histogram equalization processing on the vehicle surface fusion image to obtain a processed image; the improved histogram equalization processing includes introducing contrast limitation and histogram redistribution in the process of calculating the cumulative distribution function; S3: Position the crack candidate area of ​​the processed image based on adaptive threshold segmentation and connected domain analysis; S4: extracting features of the located crack candidate regions using an improved ResNet deep learning network to construct a set of crack feature vectors; the improved ResNet deep learning network includes introducing a channel and spatial attention module to adaptively focus on the key regions and important feature channels of the cracks; S5: A multi-scale fusion algorithm based on the attention mechanism is used to classify the crack feature vector and output the crack location and category results.

2. The method for detecting cracks in rail vehicles based on machine vision according to claim 1, characterized in that: The step S1 comprises the following steps: S11: Use a high-speed camera to collect multiple frames of continuous images of the surface of the rail vehicle, and obtain an image sequence I = {I1,…,I t ,…,I N }, where I t is the t-th frame image, t is the frame number, N is the total number of frames in the image sequence, and the value range is t∈[1,N]; S12: Perform motion estimation on adjacent frames in the image sequence and calculate the motion vector field M between adjacent frames. t : Among them, M t (x, y) is the motion vector of the t-th frame image at the pixel coordinate (x, y), dx and dy are the horizontal and vertical displacements, w is the search window size, i and j are the horizontal and vertical offsets within the search window, I t (x+i, y+j) is the pixel value of the t-th frame image at the pixel coordinate (x+i, y+j), I t-1 (x+dx+i,y+dy+j) is the pixel value of the t-1th frame image at the pixel coordinate (x+dx+i,y+dy+j), ‖·‖2 is the L2 norm; S13: Perform multi-frame image fusion based on the motion vector field to obtain a compensated vehicle surface fusion image I fusion .

3. The method for detecting cracks in rail vehicles based on machine vision according to claim 2, characterized in that: The step S13 comprises the following steps: S131: Perform spatial transformation on each frame of image according to the motion vector field to obtain a compensated single frame image: in, The image after compensation for frame t The pixel value at pixel coordinates (x,y), is the pixel value of the original image of the tth frame at the coordinate after displacement, and They are the motion vector field M t The horizontal and vertical components at the pixel coordinate (x,y), This is a round down operation; S132: Perform multi-frame image fusion on the compensated single-frame image to obtain a vehicle surface fusion image: Among them, I fusion (x, y) is the vehicle surface fusion image I fusion The pixel value at pixel coordinates (x, y), W t (x,y) is the compensated image of the tth frame The weight coefficient at the pixel coordinate (x,y), The image after compensation for frame t The gradient value at the pixel coordinate (x, y), σ is the weight decay speed control parameter, and e is a natural constant.

4. The method for detecting cracks in rail vehicles based on machine vision according to claim 3, characterized in that: The step S2 comprises the following steps: S21: Adaptive local contrast enhancement of vehicle surface fusion image: I contrast (x,y)=μ local (x,y)+β(x,y)×(I fusion (x,y)-μ local (x,y)) Among them, I contrast (x,y) is the enhanced image I contrast The pixel value at pixel coordinate (x,y), μ local (x, y) is the vehicle surface fusion image I fusion The pixel mean in the 5×5 local area centered at the pixel coordinate (x, y), β(x, y) is the adaptive gain coefficient, γ is the global gain coefficient, σ global The fused image I is the vehicle surface fusion The global pixel standard deviation, σ local (x, y) is the vehicle surface fusion image I fusion The standard deviation of pixels within a 5×5 local area centered at the pixel coordinate (x,y); S22: for the enhanced image I contrast Perform improved histogram equalization processing.

5. The method for detecting cracks in rail vehicles based on machine vision according to claim 4, characterized in that: The step S22 comprises the following steps: S221: Calculate the cumulative distribution function: Where level is the grayscale value, level∈[0,L-1], L is the grayscale level, and h(k,x,y) is the number of pixels with grayscale value k in a 5×5 local area centered on the pixel coordinate (x,y); S222: Contrast Limitation and Histogram Redistribution: in, is the histogram value after limitation, α is the contrast limitation coefficient, and E(x,y) is the number of pixels exceeding the limit; Evenly distribute the excess pixels to all gray levels and get the reallocated histogram: S223: Obtain the processed image: Update the cumulative distribution function: Use the updated cumulative distribution function to calculate and obtain the processed image: I process (x,y)=(L-1)×CDF new (I contrast (x,y),x,y) Among them, I process (x,y) is the processed image I process The pixel value at pixel coordinates (x,y).

6. The method for detecting cracks in rail vehicles based on machine vision according to claim 5, characterized in that: The step S3 comprises the following steps: S31: Calculate the local adaptive threshold: Where T(x,y) is the local adaptive threshold at the pixel coordinate (x,y), is the processed image I process The pixel mean in a 7×7 local area centered at the pixel coordinate (x, y), where C is a fixed offset; S32: Binarization of the processed image: Among them, I binary (x,y) is the binary image I binary The pixel value at pixel coordinate (x,y); S33: Morphological processing of binary images: I morph =Close(Open(I binary ,K1),K2) Among them, I morph is the result of morphological processing, Open(·) and Close(·) are the opening and closing operations respectively, K1 and K2 are the morphological operator kernels; S34: Connected domain analysis and geometric feature screening: Among them, {R p′ } is the connected domain set, N′ is the number of connected domains, ConnectedComponents(·) is the connected domain analysis function, λ p′ and ω p′ are the length and width of the p′th connected domain, A p is the area of ​​the p′th connected domain, τ1 is the aspect ratio threshold, τ2 is the area threshold, R crack is the set of crack connected domains after screening; S35: Calculate the minimum enclosing rectangle as the crack candidate area: in, and are the horizontal and vertical coordinates of the center of the candidate region, respectively. and are the width and height of the candidate region, respectively; MinBoundingRect(·) is the minimum bounding rectangle calculation function; R p is the pth crack connected domain.

7. The method for detecting cracks in rail vehicles based on machine vision according to claim 6, characterized in that: The step S4 comprises the following steps: S41: Normalize the size of each crack candidate region: in, is the pth normalized crack candidate region image, Crop(I process ,B p ) is the processed image I process Cut out the B p For the boundary area, To resize the input image to a standard size, a bilinear interpolation operation is performed. and are the standardized width and height respectively; S42: Building an improved ResNet deep learning network: Among them, F1, F2 and F3 are the feature maps obtained after their respective operations, Conv(·,θ1) is the convolution operation function based on parameter θ1, ResBlock(·,θ2) is the residual block operation function based on parameter θ2, and CBAM(·,θ3) is the channel and spatial attention module based on parameter θ3. is the number of channels of feature map F1, is the number of channels of feature maps F2 and F3, is the height of feature map F1, is the height of feature maps F2 and F3, is the width of feature map F1, is the width of feature maps F2 and F3, is the field of real numbers; S43: Extract feature vectors and construct crack feature vector sets: V={v p |R p ∈R crack } Among them, GAP(·) is the global average pooling operation function, v p is the feature vector of the pth crack candidate region, and V is the crack feature vector set consisting of the feature vectors of all crack candidate regions.

8. The method for detecting cracks in rail vehicles based on machine vision according to claim 7, characterized in that: The step S5 comprises the following steps: S51: Constructing a multi-scale feature pyramid: Q i′ =DownSample(V,ε i′ ),i′∈{1,2,3} Among them, DownSample(·,ε i′ ) is the downsampling scale factor ε based on the i′th i′ Function to downsample feature vectors, UpSample(·,δ i′ ) is the upsampling scale factor δ based on the i′th i′ The upsampling function, Q i′ is the feature of the i′th layer feature pyramid, Q i′ Features after upsampling to a unified scale; S52: Calculate attention weight: Among them, SoftAttention(·,ψ) is the function that calculates the soft attention weight based on the parameter ψ, φ i′ is the soft attention weight of the i′th layer feature; S53: Feature Fusion and Classification: Among them, Z is the fused feature vector, Based on parameters Function for fully connected layer calculation, Softmax(·) is the normalization function, Prob(cls|Z) is the category probability distribution under the given feature Z, cls∈{0,1,…,η} is the crack category, η is the total number of crack categories; S54: Crack detection result output: G={(B p ,cls * )} Among them, cls * is the predicted crack category, and G is the final crack detection result set, including crack location and crack category.

9. A rail vehicle crack detection system based on machine vision, characterized in that: include: Image preprocessing module: obtain continuous images of the surface of the moving rail vehicle and perform motion blur compensation, fuse the compensated images to obtain a fused image of the vehicle surface; Image enhancement module: performs adaptive contrast enhancement and improved histogram equalization processing on the vehicle surface fusion image to obtain the processed image; Crack candidate region positioning module: locates crack candidate regions in processed images based on adaptive threshold segmentation and connected domain analysis; Feature extraction module: Use the improved ResNet deep learning network to extract features from the located crack candidate areas and construct a set of crack feature vectors; Classification and output module: A multi-scale fusion algorithm based on the attention mechanism is used to classify the crack feature vector and output the crack location and category results; To realize the rail vehicle crack detection method based on machine vision as described in any one of claims 1-8.

Citation Information

Patent Citations

  • A device crack detection system based on millimeter wave imaging radar

    CN119310110B

Cited By

  • Aircraft engine part fatigue crack propagation identification method based on machine vision

    CN120509325A

  • Railway track crack defect detection method and device based on machine vision

    CN121366160A