Railway wheel surface defect intelligent identification method based on deep learning

By combining polar coordinate expansion and deep learning modeling, a circumferential geometric position encoding, morphological prior embedding, and adaptive token routing mechanism were designed to solve the problem of unstable identification of railway wheel surface defects under complex working conditions, and to achieve high-precision and robust defect identification.

CN120997194APending Publication Date: 2025-11-21SHANXI SHENGHENGYUAN MACHINERY PROCESSING CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511276478.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing methods for detecting defects on the surface of railway wheels are susceptible to interference from lighting, reflection, and surface contamination under complex working conditions, resulting in insufficient recognition accuracy. They also struggle to fully utilize the circumferential geometric features of the wheel surface, and have limitations in multi-scale feature fusion and contextual modeling, leading to unstable recognition.

Method used

By employing a deep learning-based approach, combining polar coordinate unfolding and deep learning modeling, a circumferential geometric position encoding, morphological prior embedding, state space hybridization, and adaptive token routing mechanism are designed. Utilizing the circumferential periodicity features of the wheel surface and fusing multi-scale contextual information, accurate identification of cracks and pitting defects is achieved.

Benefits of technology

It maintains recognition accuracy under complex lighting and reflective interference conditions, improves the integrity and stability of defect edge recognition, and achieves accurate and reliable recognition of defect location, contour, category and length, with high accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997194A_ABST
    Figure CN120997194A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning-based railway wheel surface defect intelligent identification method, which comprises the following steps of: obtaining an original image of a railway wheel surface and carrying out image preprocessing to generate a standardized polar coordinate image set and a preprocessing parameter set; an enhanced DINOv2 model is constructed; calculating and generating a circumferential geometric coding feature tensor according to the preprocessing parameter set; according to the standardized polar coordinate image set and the circumferential geometric coding feature tensor, fusion is carried out to generate a fusion feature tensor; carrying out annular long dependence modeling and sequence mixing processing on the fusion feature tensor; generating a token routing index according to the multi-source features, and extracting a high-resolution token set and a low-resolution token set; and executing cross-scale decoding, and outputting a railway wheel surface defect identification result. According to the method, the crack pitting detection rate is increased, the edge contour precision is improved, and the recognition stability under complex working conditions is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of railway vehicle inspection, and in particular to a method for intelligent identification of surface defects on railway wheels based on deep learning. Background Technology

[0002] As a critical load-bearing component of trains, the surface defects of railway wheels directly affect operational safety and lifespan. Existing inspection methods mostly rely on manual visual inspection or traditional image processing methods. These methods are easily affected by lighting, reflection, and surface contamination when dealing with early, minute defects such as cracks and pitting under complex working conditions, resulting in insufficient identification accuracy.

[0003] In recent years, deep learning has been applied in industrial defect identification, with some methods attempting to detect wheel surface defects using convolutional neural networks or transfer learning models. However, these methods typically rely on fixed spatial feature extraction mechanisms, making it difficult to fully utilize the circumferential geometric features of the wheel surface. In polar coordinate unfolded images, they are prone to boundary discontinuities and angular feature distortion, thus affecting the model's ability to characterize defect edges and geometric details.

[0004] Furthermore, existing technologies have limitations in multi-scale feature fusion and contextual modeling. Common feature fusion methods cannot simultaneously consider the global structure and local details of defects, leading to instability in the model when identifying key indicators such as crack length and defect contour. In existing detection frameworks, the allocation of tokens or features is mostly based on fixed strategies, lacking the ability to adaptively schedule features at different resolutions, which limits the recognition accuracy and robustness in complex scenarios.

[0005] Therefore, how to provide a deep learning-based intelligent identification method for railway wheel surface defects is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an intelligent identification method for railway wheel surface defects based on deep learning. This invention combines polar coordinate unfolding and deep learning modeling, and designs processing mechanisms such as circumferential geometric position encoding, morphological prior embedding, state space mixing, and adaptive token routing to achieve accurate identification of cracks and pitting defects on railway wheel surfaces. This method can fully utilize the circumferential periodic features of the wheel surface and integrate multi-scale contextual information, maintaining identification accuracy even under complex lighting and reflective interference conditions. It possesses advantages such as high accuracy, strong robustness, and good adaptability.

[0007] A method for intelligent identification of railway wheel surface defects based on deep learning according to an embodiment of the present invention includes the following steps:

[0008] The original image of the railway wheel surface is acquired and preprocessed to generate a set of standardized polar coordinate images and a set of preprocessing parameters.

[0009] An enhanced DINOv2 model is constructed, which includes a circumferential geometric position encoding module, a morphological prior embedding module, a state space mixing module, a defect adaptive token routing module, and a segmentation detection decoding module.

[0010] The circumferential geometric position encoding module calculates the circumferential relative angle difference of pixel pairs in the polar coordinate domain based on the preprocessed parameter set, and performs vectorized encoding under circumferential periodic constraints to generate circumferential geometric encoding feature tensors.

[0011] In the morphological prior embedding module, morphological operations based on parameterized structuring elements are performed on a set of standardized polar coordinate images to generate a morphological prior channel tensor, which is then fused with the circumferential geometric coding feature tensor at the channel level to obtain a fused feature tensor.

[0012] Based on the fusion feature tensor, circumferential long dependency modeling and sequence mixing processing are performed in the state space mixing module to generate a mixed backbone feature set;

[0013] In the defect adaptive token routing module, a token routing index is generated based on the hybrid backbone feature set, circumferential geometric coding feature tensor, morphological prior channel tensor and preprocessing parameter set, and the high-resolution token set and low-resolution token set are extracted.

[0014] In the segmentation detection and decoding module, cross-scale decoding is performed based on the hybrid backbone feature set, combined with the high-resolution token set and the low-resolution token set, to output the identification results of railway wheel surface defects.

[0015] Optionally, the generation of the standardized polar coordinate image set and the preprocessing parameter set includes:

[0016] The image preprocessing includes locating the wheel flange of the original image of the railway wheel surface, mapping the wheel flange region into a polar coordinate unfolded image, and performing circumferential periodic filling, illumination normalization, reflection suppression and small-angle rotation enhancement on the polar coordinate unfolded image to generate a standardized set of polar coordinate images.

[0017] In the image preprocessing stage, each processing step generates corresponding parameters, which together constitute a preprocessing parameter set. The preprocessing parameter set includes rim positioning parameters, polar coordinate unfolding parameters, circumferential periodic filling parameters, illumination normalization parameters, reflection suppression parameters, and rotation enhancement parameters.

[0018] Optionally, the generation of the circumferential geometric coding feature tensor includes:

[0019] The standardized polar coordinate image set and the preprocessing parameter set are input into the circumferential geometric position encoding module to determine the angular and radial coordinates for each pixel in the polar coordinate domain.

[0020] Based on the rotation enhancement parameters in the preprocessing parameter set, the angular coordinate difference between any pair of pixels is calculated, and the calculated angular coordinate difference is limited to the rotation angle range defined by the rotation enhancement parameters to obtain the constrained angular difference.

[0021] In the process of calculating the angular difference, a circumferential periodic constraint is introduced, and the zero angle point and the maximum angle point of the polar coordinate domain boundary are treated as continuous boundaries.

[0022] Based on the angular difference after being processed by limiting the rotation angle range and circumferential periodic constraints, a vectorization encoding operation is performed to obtain an angular position encoding representation that matches the polar coordinate domain;

[0023] In the local neighborhood of each pixel, the diagonal positional encoding representation is spatially converged and aligned with the standardized polar coordinate image set in spatial position to form a circumferential geometric encoding feature map.

[0024] Based on the circumferential geometric coding feature map, the circumferential geometric coding feature tensor is generated using the ViT Patch Embedding method.

[0025] Optionally, the generation of the fusion feature tensor includes:

[0026] In the polar coordinate domain, a two-dimensional Gaussian function is used to parameterize the convolution kernel in the radial and angular directions to generate a set of parameterized structuring elements. The set of parameterized structuring elements is updated during training through the backpropagation algorithm.

[0027] In the morphological prior embedding module, based on the parameterized structuring element set, dilation and erosion operations are performed on the standardized polar coordinate image set to obtain dilation response map and erosion response map, which are then combined sequentially to form opening operation response map and closing operation response map.

[0028] Channel stacking and angular consistency convergence are performed on the dilation response map, corrosion response map, opening operation response map and closing operation response map to generate morphological a priori channel tensor.

[0029] In the channel dimension, the morphological prior channel tensor and the circumdirectional geometric coding feature tensor are concatenated, and channel-level fusion and intensity normalization are performed to obtain the fused feature tensor.

[0030] Optionally, the calculation methods for the expansion and erosion operations are defined as follows:

[0031]

[0032] M(p i )=D(p i )-E(p i );

[0033] Wherein, D(p) i E(p) represents the expansion response function. i ) represents the corrosion response function, I(·) represents the grayscale function of the normalized polar coordinate image set, p i represents the pixel position index in polar coordinates, and b represents the local displacement vector of the structuring element. r Let b represent the radial component of the displacement vector b. θ This represents the angular component of the displacement vector b. The support set representing the parameterized structuring element is defined by the radial scale parameter σ. r Angular scale parameter σ θ Together with the shape parameter κ, w b This represents the weight parameter corresponding to the parameterized structuring element set at displacement position b. Represents the Gaussian decay function. p represents the angular periodicity correction term. θ M(p) represents the angular periodicity parameter. i ) represents the expansion response D(p) i Corrosion response E(p) i The morphological prior channel response is obtained by the difference of ).

[0034] Optionally, the generation of the high-resolution token set and the low-resolution token set includes:

[0035] The hybrid backbone feature set, circumferential geometric coding feature tensor, morphological prior channel tensor and preprocessing parameter set are input into the defect adaptive token routing module, which performs concatenation in the channel dimension and generates a candidate feature matrix through a linear transformation.

[0036] The candidate feature matrix is ​​scored using a circumferential weighted method. The weight vector is jointly determined by the angular position vector of the circumferential geometrically encoded feature tensor and the edge strength of the morphological prior channel tensor, generating a score value for each candidate feature.

[0037] The score values ​​are normalized and mapped to the rotation enhancement parameters in the preprocessing parameter set to obtain a distribution vector with a consistent numerical range, and a token routing index is generated based on the distribution vector.

[0038] The candidate feature matrix is ​​allocated according to the token routing index. Features with scores higher than the threshold are selected to form a high-resolution token set, and features with scores lower than the threshold are selected to form a low-resolution token set.

[0039] Preserve the original spatial resolution in the high-resolution token set, and perform downsampling in the low-resolution token set.

[0040] Optionally, the generation of the railway wheel surface defect identification results includes:

[0041] The hybrid backbone feature set and the high-resolution token set are aligned in the spatial dimension, and the features are fused through a cross-scale attention mechanism. The feature fusion result is then concatenated with the low-resolution token set in the channel dimension, and global context information is extracted through a multi-layer convolutional decoder to obtain cross-scale fused features.

[0042] A pixel-level segmentation mask is generated on the cross-scale fusion features, and the defect contour of the railway wheel surface is extracted using the pixel-level segmentation mask.

[0043] Bounding box regression is performed on the boundary of the defect contour to obtain the defect location, and the arc length of the defect along the rim direction is calculated based on the morphological skeleton within the region of the pixel-level segmentation mask as the defect length.

[0044] In the classification prediction branch based on cross-scale fusion feature settings, class prediction is performed on the defect region, and the defect category is output.

[0045] The railway wheel surface defect identification results are generated, including the defect location, defect outline, defect category and defect length.

[0046] The beneficial effects of this invention are:

[0047] First, this invention can perform feature modeling based on the circumferential periodicity of the polar coordinate unfolded image of the railway wheel surface, avoiding the problem of feature discontinuity at the angular boundary in existing methods, thereby improving the integrity and stability of defect edge recognition.

[0048] Secondly, by introducing a morphological prior embedding mechanism of parameterized structural elements, this invention achieves better results in enhancing the edge features of minute defects such as cracks and pitting, effectively overcoming the shortcomings of traditional methods in terms of insufficient recognition accuracy under conditions of light interference and surface reflection.

[0049] Furthermore, this invention utilizes the collaborative modeling capabilities of state-space mixing and adaptive token routing to achieve the fusion of multi-scale contextual information and local details, ensuring that the identification results of defect location, contour, category, and length are more accurate and reliable, and that the overall invention possesses high accuracy, robustness, and adaptability. Attached Figure Description

[0050] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0051] Figure 1 This is an overall flowchart of a deep learning-based intelligent identification method for railway wheel surface defects proposed in this invention.

[0052] Figure 2 This is a schematic diagram of the circumferential geometric position encoding module in this invention;

[0053] Figure 3 This is a schematic diagram of the morphological prior embedding module in this invention. Detailed Implementation

[0054] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0055] refer to Figure 1-3 A method for intelligent identification of surface defects in railway wheels based on deep learning includes the following steps:

[0056] The original image of the railway wheel surface is acquired and image preprocessing is performed. The image preprocessing includes wheel flange positioning, mapping the wheel flange area into a polar coordinate unfolded image, performing circumferential periodic filling, illumination normalization, reflection suppression and small-angle rotation enhancement, and generating a standardized polar coordinate image set and a preprocessing parameter set.

[0057] An enhanced DINOv2 model is constructed, which includes a circumferential geometric position encoding module, a morphological prior embedding module, a state space mixing module, a defect adaptive token routing module, and a segmentation detection decoding module.

[0058] The standardized polar coordinate image set and the preprocessing parameter set are input into the circumferential geometric position encoding module. The circumferential geometric position encoding module calculates the circumferential relative angle difference of pixel pairs in the polar coordinate domain according to the rotation enhancement parameters limited by the preprocessing parameter set, and performs vectorization encoding under circumferential periodic constraints to generate a circumferential geometric encoding feature tensor. The circumferential geometric encoding feature tensor and the standardized polar coordinate image set are then simultaneously sent into the morphological prior embedding module.

[0059] In the morphological prior embedding module, morphological operations based on parameterized structuring elements are performed on a set of standardized polar coordinate images to generate a morphological prior channel tensor. The parameters of the parameterized structuring elements are updated during training to highlight the edge features of cracks and pitting. The fused feature tensor is obtained by channel-level fusion with the circumferential geometric coding feature tensor and then input into the state space mixing module.

[0060] Based on the fusion feature tensor, circumferential long dependency modeling and sequence mixing processing are performed in the state space mixing module to generate a mixed backbone feature set, and the mixed backbone feature set is input into the defect adaptive token routing module.

[0061] In the defect adaptive token routing module, a token routing index is generated based on the hybrid backbone feature set, circumferential geometric coding feature tensor, morphological prior channel tensor and preprocessing parameter set. The high-resolution token set and the low-resolution token set are extracted and input into the segmentation detection decoding module.

[0062] In the segmentation detection and decoding module, based on the hybrid backbone feature set, combined with the high-resolution token set and the low-resolution token set, cross-scale decoding is performed to output the defect identification results of the railway wheel surface. The defect identification results include the defect location, defect contour, defect category and defect length.

[0063] The hybrid backbone feature set, high-resolution token set, and low-resolution token set are defined as multi-source features.

[0064] This invention proposes a deep learning-based intelligent identification method for railway wheel surface defects. It constructs a standardized polar coordinate image set and a preprocessing parameter set through image preprocessing, and sequentially introduces circumferential geometric position encoding, morphological prior embedding, state-space mixing, adaptive token routing, and cross-scale decoding into the enhanced DINOv2 model to achieve high-precision defect identification. This method overcomes the limitations of traditional methods in boundary continuity and multi-scale feature fusion by combining polar coordinate unfolding with circumferential periodic modeling.

[0065] In this embodiment, the generation of the standardized polar coordinate image set and the preprocessing parameter set includes:

[0066] The image preprocessing includes locating the wheel flange of the original image of the railway wheel surface, mapping the wheel flange region into a polar coordinate unfolded image, and performing circumferential periodic filling, illumination normalization, reflection suppression and small-angle rotation enhancement on the polar coordinate unfolded image to generate a standardized set of polar coordinate images.

[0067] In the image preprocessing stage, the original image of the railway wheel surface undergoes operations such as wheel flange positioning, polar coordinate unfolding, circumferential periodic filling, illumination normalization, reflection suppression, and small-angle rotation enhancement. Each processing step generates corresponding parameters, which together constitute a preprocessing parameter set. The preprocessing parameter set includes wheel flange positioning parameters, polar coordinate unfolding parameters, circumferential periodic filling parameters, illumination normalization parameters, reflection suppression parameters, and rotation enhancement parameters.

[0068] Specifically, the set of preprocessing parameters includes the following:

[0069] Wheel rim localization parameters: When locating the wheel rim region in the original image, the generated wheel rim center coordinates, radius estimate, and boundary offset correction amount are used. Polar coordinate unfolding parameters: When mapping the wheel rim region to a polar coordinate unfolded image, the generated radial sampling step size, angular sampling resolution, and mapping range are used. Circumferential periodic filling parameters: When performing circumferential stitching at the boundary of the polar coordinate unfolded image, the generated boundary alignment offset and filling step size are used. Illumination normalization parameters: During image grayscale normalization, the generated normalized mean, normalized variance, and illumination smoothing coefficient are used. Reflection suppression parameters: When suppressing local strong reflective areas, the generated reflection detection threshold, suppression weight factor, and repair area mask are used. Rotation enhancement parameters: When performing small-angle rotation enhancement, the generated upper limit of the rotation angle range, rotation step size, and corresponding rotation direction indication are used.

[0070] The preprocessing parameter set is used not only to record the operation information of the image preprocessing stage in subsequent steps, but also as input conditions to participate in the circumferential geometric position encoding module and the defect adaptive token routing module, so as to ensure that the circumferential feature modeling and token allocation process can be constrained and corrected according to the rotation angle range and boundary conditions of the wheel surface image.

[0071] In this embodiment, the generation of the circumferential geometric coding feature tensor includes:

[0072] The standardized polar coordinate image set and the preprocessing parameter set are input into the circumferential geometric position encoding module to determine the angular and radial coordinates for each pixel in the polar coordinate domain.

[0073] Based on the rotation enhancement parameters in the preprocessing parameter set, the angular coordinate difference between any pair of pixels is calculated, and the calculated angular coordinate difference is limited to the rotation angle range defined by the rotation enhancement parameters to obtain the constrained angular difference.

[0074] In the process of calculating the angular difference, a circumferential periodic constraint is introduced, and the zero angle point and the maximum angle point of the polar coordinate domain boundary are treated as continuous boundaries to ensure the continuity of the angular difference in the circumferential dimension.

[0075] Based on the angular difference after being processed by limiting the rotation angle range and circumferential periodic constraints, a vectorization encoding operation is performed to obtain an angular position encoding representation that matches the polar coordinate domain;

[0076] In the vectorized coding operation, based on the corrected angular difference, a two-dimensional coding vector is constructed using sine and cosine functions, so that the relative position information of each pixel pair in the angular dimension is expressed in the form of a periodic function, resulting in an angular position coding representation that matches the polar coordinate domain, thereby ensuring the continuity of features in the wheel circumferential direction and the consistency of rotation enhancement.

[0077] In the local neighborhood of each pixel, the diagonal positional encoding representation is spatially converged and aligned with the standardized polar coordinate image set in spatial position to form a circumferential geometric encoding feature map.

[0078] Based on the circumdirectional geometric coding feature map, a circumdirectional geometric coding feature tensor is generated using the ViT Patch Embedding method. The standardized polar coordinate image set and the circumdirectional geometric coding feature tensor are then input into the morphological prior embedding module.

[0079] In this embodiment, the generation of the fusion feature tensor includes:

[0080] In the polar coordinate domain, a two-dimensional Gaussian function is used to parameterize the convolution kernel in the radial and angular directions to generate a set of parameterized structural elements. The parameterized structural elements are defined by radial scale parameters, angular scale parameters, directional parameters, and shape parameters. The radial scale parameter controls the expansion range of the structural element in the radial direction, the angular scale parameter controls the coverage range of the structural element in the angular direction, the directional parameter adjusts the rotation direction of the structural element in the polar coordinate domain, and the shape parameter limits the ellipticity or aspect ratio of the structural element. The set of parameterized structural elements is updated through the backpropagation algorithm during training, so that the structural elements can adaptively fit the slender shape of the crack and the local contour of the pitting in different regions of the wheel surface image, thereby highlighting the edge features of defects on the railway wheel surface in subsequent morphological operations.

[0081] In the morphological prior embedding module, based on the parameterized structuring element set, dilation and erosion operations are performed on the standardized polar coordinate image set to obtain dilation response map and erosion response map, which are then combined sequentially to form opening operation response map and closing operation response map.

[0082] Channel stacking and angular consistency convergence are performed on the dilatation response map, corrosion response map, opening operation response map, and closing operation response map to generate a morphological prior channel tensor. The angular consistency convergence is a feature fusion method designed to address the problem of periodic boundaries in the angular dimension of the polar coordinate unfolded image of the railway wheel surface. After channel stacking of the feature maps, the angular consistency convergence periodically aligns the features at the positions of θ=0 and θ=2π and performs a weighted average in the circumferential dimension, so that the features at the boundary can maintain continuity. This ensures that the edge features of defects such as cracks and pitting in the circumferential direction in the polar coordinate domain can be completely captured and transmitted to the morphological prior channel tensor.

[0083] In the channel dimension, the morphological prior channel tensor and the circumferential geometric coding feature tensor are concatenated, and channel-level fusion and intensity normalization are performed to obtain the fused feature tensor.

[0084] The output fused feature tensor is used as input to the state space mixing module.

[0085] This step proposes a morphological prior embedding method based on parameterized structuring elements. In the polar coordinate domain, a set of trainable structuring elements is generated using a Gaussian function. Dilation, erosion, and opening / closing operations are performed on the image. A morphological prior channel tensor is generated by combining channel stacking and angular consistency convergence. This tensor is then fused with a circumferential geometric coding feature tensor to obtain a fused feature tensor, enabling morphological operations to dynamically adapt to defect edge features.

[0086] In this embodiment, the calculation methods for the expansion and corrosion operations are defined as follows:

[0087]

[0088] M(p i )=D(p i )-E(p i );

[0089] Wherein, D(p) i ) represents the dilation response function for pixel p. i The result after parameterized dilation calculation, E(p) i ) represents the corrosion response function for pixel p. i The result after parametric erosion operation, I(·) represents the grayscale function of the standardized polar coordinate image set, p i Let represent the pixel position index in polar coordinates, represent the i-th pixel, and represent the local displacement vector of the structuring element, containing radial and angular components. r The radial component of the displacement vector b represents the offset in the radial direction. θ Let represent the angular component of the displacement vector b, indicating the offset in the angular direction. The support set representing the parameterized structuring element is defined by the radial scale parameter σ. r Angular scale parameter σ θ Together with the shape parameter κ, the radial scale parameter controls the range of expansion or contraction in the radial direction, the angular scale parameter controls the range of expansion or contraction in the angular direction, and the shape parameter describes the morphological characteristics of the structural element (such as ellipse, strip, etc.). b This represents the weight parameters corresponding to the parameterized structuring element set at displacement position b. These parameters are updated during training to adjust the influence strength of the structuring elements. This represents the Gaussian decay function, used to impose scale constraints on radial and angular offsets, ensuring that the influence away from the center is reduced. This represents the angular periodicity correction term, used to introduce angular periodicity to ensure boundary continuity. θ M(p) represents the angular periodicity parameter, and M(p) represents the complete angular period length. i ) represents the expansion response D(p) i Corrosion response E(p) i The morphological prior channel response obtained by the difference of ) is used to highlight the edge features of defects on the surface of railway wheels.

[0090] This step defines the calculation method for dilation and corrosion operations by combining parameterized structuring elements with Gaussian decay and angular periodic correction. It also forms a morphological prior channel response through difference, transforming the traditional morphological operation formula into an optimizable function framework, enabling it to have adaptive learning capabilities and enhancing the model's ability to express crack and pitting edges.

[0091] In this embodiment, the generation of the high-resolution token set and the low-resolution token set includes:

[0092] The hybrid backbone feature set, circumferential geometric coding feature tensor, morphological prior channel tensor and preprocessing parameter set are input into the defect adaptive token routing module, which performs concatenation in the channel dimension and generates a candidate feature matrix through a linear transformation.

[0093] The candidate feature matrix is ​​scored using a circumferential weighted method. The weight vector is jointly determined by the angular position vector of the circumferential geometrically encoded feature tensor and the edge strength of the morphological prior channel tensor, generating a score value for each candidate feature.

[0094] When calculating the weights, the angular position vector corresponding to each pixel is first extracted from the circumferential geometric coding feature tensor. This vector represents the angular position information of the pixel in the polar coordinate domain. At the same time, the edge intensity value of the corresponding pixel is extracted from the morphological prior channel tensor. This value reflects the salience of the defect boundary. Then, the angular position vector and the edge intensity value are normalized to make them have a consistent numerical range. Finally, the two are combined by linear combination to generate the weight vector.

[0095] After obtaining the weight vector, it is element-wise multiplied with the candidate feature matrix to weight the candidate features along the channel dimension. A fully connected layer is then used for feature compression, yielding a scalar response value as the initial score for the candidate feature. To ensure the scores are comparable within a uniform range, the scores of all candidate features are normalized using the Softmax function to transform the scores into a probability distribution. The final probability value is the score for each candidate feature; a higher value indicates greater importance of the feature in the defect identification process, and it is preferentially assigned to the high-resolution token set.

[0096] The score values ​​are normalized and mapped to the rotation enhancement parameters in the preprocessing parameter set to obtain a distribution vector with a consistent numerical range. The token routing index is then generated based on the distribution vector. The token routing index is used to characterize the allocation priority of features in high-resolution and low-resolution paths.

[0097] When generating the token routing index, the score value of each candidate feature is first normalized according to the rotation enhancement parameters in the preprocessing parameter set. That is, the rotation angle range defined by the rotation enhancement parameters is used as the normalization interval, mapping all score values ​​to a unified numerical range of [0,1]. Then, a thresholding operation is performed on the normalized distribution vector. Features with values ​​greater than the threshold θ are marked as high-resolution paths, and features with values ​​less than or equal to the threshold θ are marked as low-resolution paths. The threshold θ can be adaptively set according to the value range of the rotation enhancement parameters. Finally, the identifier number of each candidate feature is combined with the corresponding path label to form the token routing index, which is used to represent the allocation priority of the feature in high-resolution paths and low-resolution paths.

[0098] The candidate feature matrix is ​​allocated according to the token routing index. Features with scores higher than the threshold are selected to form a high-resolution token set, and features with scores lower than the threshold are selected to form a low-resolution token set.

[0099] The original spatial resolution is maintained in the high-resolution token set, a downsampling operation is performed in the low-resolution token set, and the high-resolution token set and the low-resolution token set are output as inputs to the segmentation detection decoding module.

[0100] This step proposes an adaptive token routing mechanism based on multi-source features. A candidate matrix is ​​generated by concatenating a hybrid backbone feature set, a circumferential geometric coding feature tensor, a morphological prior channel tensor, and a preprocessing parameter set. A token routing index is generated using joint scoring, and high- and low-resolution token sets are extracted based on a threshold. This breaks through the traditional fixed token allocation method, enabling high- and low-resolution features to be adaptively and collaboratively processed, thus improving cross-scale modeling capabilities.

[0101] In this embodiment, the generation of the railway wheel surface defect identification result includes:

[0102] The hybrid backbone feature set and the high-resolution token set are aligned in the spatial dimension, and the feature is fused through a cross-scale attention mechanism to preserve the spatial details of the defect edges. The feature fusion result is then concatenated with the low-resolution token set in the channel dimension, and global context information is extracted through a multi-layer convolutional decoder to enhance the overall structural consistency of the defect region, thus obtaining cross-scale fused features.

[0103] A pixel-level segmentation mask is generated on the cross-scale fusion features, and the defect contour of the railway wheel surface is extracted using the pixel-level segmentation mask.

[0104] Bounding box regression is performed on the boundary of the defect contour to obtain the defect location, and the arc length of the defect along the rim direction is calculated based on the morphological skeleton within the region of the pixel-level segmentation mask as the defect length.

[0105] In the classification prediction branch based on cross-scale fusion feature settings, class prediction is performed on the defect region, and the defect category is output.

[0106] The railway wheel surface defect identification results are generated, including the defect location, defect outline, defect category and defect length.

[0107] Example 1:

[0108] To verify the feasibility of this invention in practice, it was applied to the defect detection task on the surface of railway vehicle wheels. The scenario involved the routine inspection of railway vehicles, requiring batch defect identification of a large number of wheels. The defect types mainly included cracks, pitting, and surface wear. Existing manual inspection and traditional convolutional neural network detection methods often suffer from high false positive rates, high false negative rates, and incomplete defect contour segmentation when faced with interference from light reflection, dirt coverage, and complex wheel flange geometry. This leads to decreased maintenance efficiency and the failure to detect some potential risks in a timely manner.

[0109] In this scenario, the method of the present invention first acquires an image of the wheel surface and performs preprocessing steps under polar coordinate unfolding conditions, including rim positioning, illumination normalization, reflection suppression, and small-angle rotation enhancement, thereby obtaining a standardized polar coordinate image with stable quality. Subsequently, an enhanced deep learning model is used to encode and decode the input image. The model design effectively utilizes the circumferential periodicity of the wheel, while combining morphological priors to enhance defect edge features. Furthermore, an adaptive token routing mechanism is used to rationally allocate features at different resolutions, resulting in a more complete defect contour and accurate location description during the segmentation and decoding process.

[0110] In practical applications, 1000 railway wheel images with crack, pitting, and wear markings were selected as a test set. The detection performance of the method of this invention was compared with two traditional schemes: an image processing method based on manual features and a detection method based on convolutional neural networks. Evaluation metrics included defect detection rate, false detection rate, contour accuracy, positional deviation, and computation time. Experimental data are shown in the table below.

[0111] Table 1. Performance comparison of different methods in railway wheel surface defect identification.

[0112]

[0113] As shown in Table 1, the method of this invention outperforms existing methods in all performance indicators. Regarding the defect detection rate, this invention achieves 91.7%, a 13.1 percentage point improvement compared to the 78.6% of the manual feature image processing method, and a 6.8 percentage point improvement compared to the 84.9% of the convolutional neural network detection method. This demonstrates better coverage of micro-cracks and pitting defects, significantly reducing missed detections. In terms of the false negative rate, this invention is only 3.3%, while the manual feature method and the convolutional neural network method are 21.4% and 11.1%, respectively, indicating a significant advantage in the comprehensiveness of defect coverage.

[0114] In terms of false detection rate, this invention achieves a false positive rate of only 3.9%, significantly lower than the 12.5% ​​of manual feature methods and the 7.8% of convolutional neural network methods. This indicates that the model can more effectively distinguish between real defects and interference information such as surface stains and reflections, reducing false alarms. Regarding contour accuracy, the segmentation IOU value of this invention is 89.4%, which is 17.1 percentage points higher than manual feature methods and 7.7 percentage points higher than convolutional neural network methods, demonstrating superior performance in defect morphology restoration, edge continuity, and contour integrity.

[0115] Regarding the positional deviation index, the error of this invention is only 0.9 mm, lower than the 2.4 mm of the manual feature method and the 1.6 mm of the convolutional neural network method, indicating more accurate defect localization and providing a more reliable reference for subsequent maintenance and processing. In terms of average processing time, this invention requires only 72 milliseconds, far lower than the 185 milliseconds of the manual feature method and the 94 milliseconds of the convolutional neural network method, demonstrating real-time detection capability and meeting the needs of rapid maintenance of railway vehicles.

[0116] The performance improvement stems from the fact that this invention combines circumferential periodic feature modeling after the wheel's polar coordinate expansion, avoiding recognition bias caused by boundary fractures; the introduction of morphological prior embedding of parameterized structural elements enhances the expressive power of crack and pitting edges; the state-space hybrid module effectively captures long-range dependencies; and the adaptive token routing mechanism achieves reasonable allocation among multi-scale features, enabling the model to consider both the global structure and local details of defects. These innovative designs collectively improve the robustness and accuracy of the model under complex working conditions, verifying the feasibility and superiority of this invention in intelligent identification of surface defects on railway wheels.

[0117] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for intelligent identification of surface defects on railway wheels based on deep learning, characterized in that, Includes the following steps: The original image of the railway wheel surface is acquired and preprocessed to generate a set of standardized polar coordinate images and a set of preprocessing parameters. An enhanced DINOv2 model is constructed, which includes a circumferential geometric position encoding module, a morphological prior embedding module, a state space mixing module, a defect adaptive token routing module, and a segmentation detection decoding module. Based on the circumferential geometric position encoding module, the circumferential relative angle difference of pixel pairs is calculated in the polar coordinate domain according to the preprocessed parameter set, and vectorized encoding is performed under the circumferential periodic constraint to generate the circumferential geometric encoding feature tensor. In the morphological prior embedding module, morphological operations based on parameterized structuring elements are performed on a set of standardized polar coordinate images to generate a morphological prior channel tensor, which is then fused with the circumferential geometric coding feature tensor at the channel level to obtain a fused feature tensor. Based on the fusion feature tensor, circumferential long dependency modeling and sequence mixing processing are performed in the state space mixing module to generate a mixed backbone feature set; In the defect adaptive token routing module, a token routing index is generated based on the hybrid backbone feature set, circumferential geometric coding feature tensor, morphological prior channel tensor and preprocessing parameter set, and the high-resolution token set and low-resolution token set are extracted. In the segmentation detection and decoding module, cross-scale decoding is performed based on the hybrid backbone feature set, combined with the high-resolution token set and the low-resolution token set, to output the identification results of railway wheel surface defects.

2. The intelligent identification method for railway wheel surface defects based on deep learning according to claim 1, characterized in that, The generation of the standardized polar coordinate image set and the preprocessing parameter set includes: The image preprocessing includes locating the wheel flange of the original image of the railway wheel surface, mapping the wheel flange region into a polar coordinate unfolded image, and performing circumferential periodic filling, illumination normalization, reflection suppression and small-angle rotation enhancement on the polar coordinate unfolded image to generate a standardized set of polar coordinate images. In the image preprocessing stage, each processing step generates corresponding parameters, which together constitute a preprocessing parameter set. The preprocessing parameter set includes rim positioning parameters, polar coordinate unfolding parameters, circumferential periodic filling parameters, illumination normalization parameters, reflection suppression parameters, and rotation enhancement parameters.

3. The intelligent identification method for railway wheel surface defects based on deep learning according to claim 1, characterized in that, The generation of the circumferential geometric coding feature tensor includes: The standardized polar coordinate image set and the preprocessing parameter set are input into the circumferential geometric position encoding module to determine the angular and radial coordinates for each pixel in the polar coordinate domain. Based on the rotation enhancement parameters in the preprocessing parameter set, the angular coordinate difference between any pair of pixels is calculated, and the calculated angular coordinate difference is limited to the rotation angle range defined by the rotation enhancement parameters to obtain the constrained angular difference. In the process of calculating the angular difference, a circumferential periodic constraint is introduced, and the zero angle point and the maximum angle point of the polar coordinate domain boundary are treated as continuous boundaries. Based on the angular difference after being processed by limiting the rotation angle range and circumferential periodic constraints, a vectorization encoding operation is performed to obtain an angular position encoding representation that matches the polar coordinate domain; In the local neighborhood of each pixel, the diagonal positional encoding representation is spatially converged and aligned with the standardized polar coordinate image set in spatial position to form a circumferential geometric encoding feature map. Based on the circumferential geometric coding feature map, the circumferential geometric coding feature tensor is generated using the ViT Patch Embedding method.

4. The intelligent identification method for railway wheel surface defects based on deep learning according to claim 1, characterized in that, The generation of the fusion feature tensor includes: In the polar coordinate domain, a two-dimensional Gaussian function is used to parameterize the convolution kernel in the radial and angular directions to generate a set of parameterized structuring elements. The set of parameterized structuring elements is updated during training through the backpropagation algorithm. In the morphological prior embedding module, based on the parameterized structuring element set, dilation and erosion operations are performed on the standardized polar coordinate image set to obtain dilation response map and erosion response map, which are then combined sequentially to form opening operation response map and closing operation response map. Channel stacking and angular consistency convergence are performed on the dilation response map, corrosion response map, opening operation response map and closing operation response map to generate morphological a priori channel tensor. In the channel dimension, the morphological prior channel tensor and the circumdirectional geometric coding feature tensor are concatenated, and channel-level fusion and intensity normalization are performed to obtain the fused feature tensor.

5. The intelligent identification method for railway wheel surface defects based on deep learning according to claim 4, characterized in that, The calculation methods for the expansion and erosion operations are defined as follows: M(p i )=D(p i )-E(p i ); Wherein, D(p) i E(p) represents the expansion response function. i ) represents the corrosion response function, I(·) represents the grayscale function of the normalized polar coordinate image set, p i represents the pixel position index in polar coordinates, and b represents the local displacement vector of the structuring element. r Let b represent the radial component of the displacement vector b. θ B(σ) represents the angular component of the displacement vector b. r ,σ θ ,κ) represents the support set of the parameterized structuring element, defined by the radial scale parameter σ. r Angular scale parameter σ θ Together with the shape parameter κ, w b This represents the weight parameter corresponding to the parameterized structuring element set at displacement position b. Represents the Gaussian decay function. p represents the angular periodicity correction term. θ M(p) represents the angular periodicity parameter. i ) represents the expansion response D(p) i Corrosion response E(p) i The morphological prior channel response is obtained by the difference of ).

6. The intelligent identification method for railway wheel surface defects based on deep learning according to claim 1, characterized in that, The generation of the high-resolution token set and the low-resolution token set includes: The hybrid backbone feature set, circumferential geometric coding feature tensor, morphological prior channel tensor and preprocessing parameter set are input into the defect adaptive token routing module, which performs concatenation in the channel dimension and generates a candidate feature matrix through a linear transformation. The candidate feature matrix is ​​scored using a circumferential weighted method. The weight vector is jointly determined by the angular position vector of the circumferential geometrically encoded feature tensor and the edge strength of the morphological prior channel tensor, generating a score value for each candidate feature. The score values ​​are normalized and mapped to the rotation enhancement parameters in the preprocessing parameter set to obtain a distribution vector with a consistent numerical range, and a token routing index is generated based on the distribution vector. The candidate feature matrix is ​​allocated according to the token routing index. Features with scores higher than the threshold are selected to form a high-resolution token set, and features with scores lower than the threshold are selected to form a low-resolution token set. Preserve the original spatial resolution in the high-resolution token set, and perform downsampling in the low-resolution token set.

7. The intelligent identification method for railway wheel surface defects based on deep learning according to claim 1, characterized in that, The generation of the railway wheel surface defect identification results includes: The hybrid backbone feature set and the high-resolution token set are aligned in the spatial dimension, and the features are fused through a cross-scale attention mechanism. The feature fusion result is then concatenated with the low-resolution token set in the channel dimension, and global context information is extracted through a multi-layer convolutional decoder to obtain cross-scale fused features. A pixel-level segmentation mask is generated on the cross-scale fusion features, and the defect contour of the railway wheel surface is extracted using the pixel-level segmentation mask. Bounding box regression is performed on the boundary of the defect contour to obtain the defect location, and the arc length of the defect along the rim direction is calculated based on the morphological skeleton within the region of the pixel-level segmentation mask as the defect length. In the classification prediction branch based on cross-scale fusion feature settings, class prediction is performed on the defect region, and the defect category is output. The railway wheel surface defect identification results are generated, including the defect location, defect outline, defect category and defect length.

Citation Information

Cited By

  • Substation metal expander top rushing detection method based on fine grit identification

    CN121564511A

  • Circuit board assembly defect feature identification method based on deep learning

    CN122265748A