Subway train brake shoe thickness state abnormity identification method based on deep learning
By using a subway inspection robot equipped with a depth camera and deep learning technology, the thickness of the thinnest part of the subway train brake shoe can be automatically identified, solving the problem of easy misjudgment by manual inspection and improving the accuracy and safety of inspection.
Patent Information
- Application Number
- CN202510878659.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-04
AI Technical Summary
During long-term operation, the brake shoes of subway trains may become thinner due to wear. Existing detection methods rely on manual inspection, which is prone to missed detections or misjudgments, affecting braking performance and safety.
A subway inspection robot carrying a depth camera was used to capture images of brake shoes. Deep learning technology was used to enhance the image quality and perform semantic segmentation to identify the thickness of the thinnest part of the brake shoe. The Hough circle transform was combined to detect the center of the wheel, calculate the brake shoe thickness, and compare it with a safety threshold.
It has achieved automated and highly stable brake shoe thickness detection, which improves the accuracy and reliability of detection, reduces human error, and ensures train safety.
Smart Images

Figure CN120894600A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent technology for rail transit equipment, specifically a method for identifying abnormal brake shoe thickness in subway trains based on deep learning. Background Technology
[0002] Urban rail transit is an important mode of transportation, and the performance of the subway train braking system directly affects the operational safety of the train and the personal safety of passengers. The brake shoes in the subway train braking system generate braking force through friction with the wheels, causing the train to slow down or stop. During long-term operation, the brake shoes wear down through continuous contact with the wheels, especially in areas where stress is concentrated or heat dissipation is uneven, which can lead to rapid thinning of the brake shoes in certain areas. When the thinnest part of the brake shoe falls below the safe limit, failure to detect and replace it in time may result in decreased braking performance or even brake failure. Therefore, during routine inspections or periodic maintenance after daily subway shutdowns, the thickness of the thinnest part of the brake shoes must be tested to ensure that it remains within the safe operating range.
[0003] Routine inspections of subway trains require maintenance personnel to possess certain professional skills, and due to the large number of inspection items, human factors can easily lead to missed inspections or misjudgments. Summary of the Invention
[0004] To address the problems of existing technologies, this invention provides a deep learning-based method for identifying abnormal brake shoe thickness in subway trains. The method utilizes a subway inspection robot to photograph the brake shoes using its own depth camera. By analyzing these images, it identifies whether the thickness of the thinnest part of the brake shoe is within the standard range. This allows the robot to automatically perform inspection tasks according to pre-defined requirements, improving the stability and reliability of the inspection process.
[0005] This invention provides a deep learning-based method for identifying anomalies in the thickness of subway train brake shoes, comprising the following steps:
[0006] Step 1) Acquire images of the target brake shoe to obtain multimodal image data, including RGB images and corresponding depth maps;
[0007] Step 2) Perform image quality enhancement processing on the multiple frames of RGB images and depth images obtained in Step 1);
[0008] Step 3) The RGB image and depth image processed in Step 2) are fed into the semantic segmentation network as inputs to perform joint segmentation inference. The gate tile semantic segmentation network based on RGBD information is adopted to fuse the texture features of the RGB image and the geometric structure information of the depth map to achieve high-precision segmentation of the gate tile region. The RGB image contains color, texture and edge cues, and each pixel value in the depth map represents the spatial distance from the point to the camera, reflecting the height undulation in the scene and strengthening the boundary between the gate tile and the surrounding environment at the structural level.
[0009] Step 4) Use Hough circle transform to find the center of the wheel in the denoised RGB image obtained in step 2; improve the accuracy of center acquisition by using prior information, which includes the wheel diameter range and the center range;
[0010] Step 5) Based on the center coordinates obtained in Step 4), obtain 360 ray equations with the center as the endpoint at equal intervals. For each ray, find the coordinates of its intersection point with the outline of the gate tile segmentation mask, and calculate the pixel length of each ray within the gate tile segmentation mask.
[0011] Step 6) Locate the thinnest part of the brake shoe, specifically:
[0012] 6.1) In step 5), each ray emanating from the center of the wheel... θ For a given thickness segment, remove all length values that are 0, indicating that the ray does not intersect with the mask;
[0013] 6.2) Sort the remaining length values by angle to form a set of thickness distributions in each direction corresponding to the center of the brake shoe:
[0014]
[0015] 6.3) In thickness sets Remove the thickness values at both ends from the set; Remove several values from the mask's edge in the top and bottom extreme directions, such as removing length values in several directions before and after, to obtain the remaining thickness value set. Calculate the minimum value: Record the coordinates P of the intersection point before and after the minimum thickness value. in and P out The line segment marked in this direction is the candidate area for the thinnest part of the brake shoe in the entire structure;
[0016] Step 7) Calculate the brake shoe thickness and compare it with the threshold to determine whether it meets the standard;
[0017] 7.1) Combining the depth map information in the RGBD image, the physical spatial distance is calculated using the three-dimensional coordinates of the two intersection points on the brake shoe contour line. The thinnest pixel thickness L obtained in step 6) is then used. minThe mapping to the actual physical thickness is as follows: Let P be the two boundary pixels at the thinnest point. in =(x in ,y in ), P out =(x out ,y out The corresponding depth value is D. in D out Using the camera intrinsic parameter matrix K, the pixels in the depth map are reconstructed into 3D point cloud points in the camera coordinate system; the corresponding intersection point point cloud 3D coordinates are (X... in ,Y in Z in ), (X out ,Y out Z out The thickness T at the thinnest point of the brake shoe real The calculation formula is as follows:
[0018]
[0019] 7.2) T real Safety threshold T for brake shoe thickness thresh The comparison is performed, and if the value is less than the threshold, it is determined that the brake shoe has reached its wear limit and needs to be replaced in time.
[0020] Further improvements include ensuring that the camera and the brake shoe surface are facing each other at a positive angle during the image acquisition process described in step 1), and continuously acquiring multiple frames of image data, that is, continuously acquiring multiple frames of RGB images and their corresponding depth images at the same viewpoint and position.
[0021] Further improvements are made to the image quality enhancement process described in step 2), which specifically involves: performing a multi-frame denoising algorithm on the RGB image to suppress random noise; and simultaneously employing temporal and spatial filtering methods on the depth image to remove discrete noise and local defects caused by sensor jitter or acquisition errors.
[0022] The multi-frame denoising algorithm specifically involves averaging the values of the same pixel across multiple frames to reduce random noise and improve the signal-to-noise ratio.
[0023] The formula for calculating the multi-frame average is as follows:
[0024] Among them: I i(x,y) This represents the pixel value of the i-th frame, where N is the frame number.
[0025] The temporal and spatial filtering methods are as follows: Assuming that n frames of depth images were acquired in step 1), they are stacked in chronological order into a three-dimensional matrix with a size of n×W×H, where n represents the number of frames in the time dimension, and W and H represent the width and height of the image, respectively; based on this three-dimensional structure, a three-dimensional median filter operator of size n×n×n is applied to each pixel in both the time and spatial dimensions to perform joint filtering processing.
[0026] The formula for calculating three-dimensional median filtering is as follows:
[0027] D' (x,y) =median(D (x+i,y+j,t+k) ), i, j, k ∈ {-2, -1, 0, 1, 2}
[0028] Where D (x,y,t) D' represents the depth value at position (x, y) in the t-th frame of the original 3D depth data. (x,y) The filtered depth value is denoted by median(·), which means taking the median value within an n×n×n neighborhood. i and j control the spatial neighborhood, and k controls the temporal neighborhood.
[0029] Further improvements, in step 3), the process of obtaining the gate shoe segmentation mask through semantic segmentation using the U-net network backbone structure is specifically as follows:
[0030] 3.1) Before being fed into the network, the RGB images are normalized and standardized, while the depth map is normalized to the [0,1] interval to ensure that the depth features are consistent under different acquisition environments;
[0031] 3.2) Design the network structure and adopt a dual-channel feature extraction strategy: RGB image and depth map are respectively input into two parallel sub-networks for preliminary feature extraction;
[0032] 3.3) The feature maps output by the two are fused along the channel dimension to make the color and geometric information fully complementary;
[0033] 3.4) The features fused in step 3.3) are fed into the U-net backbone network for further processing, and finally a gate shoe segmentation mask is generated for subsequent analysis.
[0034] Further improvements, in step 4), the Hough circle transformation algorithm specifically involves: representing the parameter space of the circle as a three-dimensional space (x... c ,y c ,r), where (x c ,y c (x) represents the coordinates of the center of the circle, and r is the radius of the circle; for edge points in the image (x) i ,y i ), satisfying the equation of a circle:
[0035] (x i -x c ) 2 +(y i -y c ) 2 =r 2
[0036] For (x) that satisfy this equation in the parameter space c ,y c We perform a voting statistics on r) and select the local maximum value as the parameter of the candidate circle.
[0037] Step 4) The method for improving the accuracy of center acquisition specifically involves: setting the radius search range of the circle [r] min ,r max The range is preset based on the structural dimensions of the wheel; the approximate center region (x0, y0) is estimated based on the vehicle structure, thereby limiting the search range of the parameter space; multiple candidate circles are evaluated based on edge closure and gradient strength, and only the optimal result is retained as the wheel center; finally, the wheel center coordinates are obtained as (x0, y0). c ,y c This point will serve as the starting point for all subsequent radiation rays.
[0038] Further improvements, step 5), specifically calculating the pixel length of each ray within the gate tile segmentation mask, involves: each ray l θ Represented as a half-line radii emanating from the center of the circle along direction θ, this ray intersects the gate segmentation mask M zero, one, or more times in the image space. A pixel-level traversal of the ray is performed to determine whether the pixels it passes through are located on the mask edge; if multiple intersection points {P1, P2, ..., P...} exist... n If the first intersection point P is retained, then the first intersection point P is retained. in Intersection with the last point P out The line segment between these two points represents the effective thickness range in that direction; the pixel length L of this line segment... θ Calculated using Euclidean distance:
[0039] The beneficial effects of this invention are as follows: by using a subway inspection robot to take pictures of the brake shoes with its own depth camera, and then analyzing these pictures, it can identify whether the thickness of the thinnest part of the brake shoe is within the standard range, and can automatically perform inspection tasks according to the set requirements, thereby improving the stability and reliability of the inspection. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating the execution process of the present invention.
[0042] Figure 2 This is a diagram of the network structure of the brake shoe segmentation network;
[0043] Figure 3 A schematic diagram of a brake shoe photograph taken by a robot;
[0044] Figure 4 This is a schematic diagram of the mask used to segment the image of the brake shoe.
[0045] Figure 5 This is a schematic diagram showing the results of center detection on the brake shoe photograph;
[0046] Figure 6 This is a schematic diagram showing the division of the brake shoe based on the center of a circle. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] This invention aims to provide a deep learning-based method for detecting the thinnest part of a brake shoe. It utilizes an RGBD camera to acquire brake shoe image data and combines it with semantic segmentation technology to accurately identify the thickness of the thinnest part of the brake shoe and determine whether it meets the standard range. First, to improve the quality of the acquired images, multiple frames of RGBD information are acquired from the brake shoe and fused. The processed RGBD image is then fed into a brake shoe semantic segmentation network to obtain a segmentation mask. Subsequently, a Hough transform is used to find the center of the wheel. Rays radiate outward from the wheel center, and the length of the line segment within the segmentation mask represents the thickness of the brake shoe at that point. The minimum value of this set of lengths is the pixel distance of the thinnest part of the brake shoe. Using the pixel coordinates and depth information at both ends of the line segment, the actual physical thickness of the thinnest part of the brake shoe can be calculated.
[0049] The technical solution adopted in this invention is implemented according to the following steps, such as... Figure 1 As shown:
[0050] Step 1: Use an RGBD camera to acquire images of the target brake shoe, obtaining multimodal image data including color images (RGB) and corresponding depth maps. The color image is shown below. Figure 3 As shown. During the acquisition process, the camera should be kept at a direct angle to the brake shoe surface, and multiple frames of image data should be acquired continuously.
[0051] Step 2: Perform image quality enhancement processing on the multi-frame RGB images and depth images obtained in Step 1: Perform multi-frame denoising algorithm on the RGB images to suppress random noise; simultaneously use temporal filtering and spatial filtering methods on the depth images to remove discrete noise and local defects caused by sensor jitter or acquisition errors.
[0052] Step 3: Input the processed RGB image and depth image from Step 2 into the semantic segmentation network and perform joint segmentation inference. The network combines the texture and color information provided by the RGB image with the spatial structure information reflected by the depth map to obtain a precise segmentation mask for the gate tile region, such as... Figure 4 As shown.
[0053] Step 4: Use Hough transform to find the center of the wheel in the denoised RGB image obtained in Step 2, such as... Figure 5 As shown, the accuracy of center acquisition can be improved by using prior information such as the wheel diameter range and the approximate center range.
[0054] Step 5: Based on the center coordinates obtained in Step 4, obtain 360 ray equations with the center as the endpoint, at 1-degree intervals. For each ray, calculate the coordinates of its intersection point with the brake shoe segmentation mask contour. A schematic diagram of the brake shoe segmentation based on the center is shown below. Figure 6 As shown, for visualization purposes, the angle intervals in this figure are 3 degrees.
[0055] Step 6: Using the intersection coordinates obtained in Step 5, calculate the pixel length of each ray segment within the gate tile segmentation mask. To improve the accuracy of finding the thinnest point of the gate tile, the five length values at the top and bottom of the segmentation mask need to be removed when finding the minimum pixel length.
[0056] Step 7: Based on the coordinates of the intersection point between the ray corresponding to the minimum pixel length obtained in Step 6 and the gate tile segmentation mask contour, calculate the actual physical distance using the depth map information. This distance is the thickness of the thinnest part of the gate tile, and compare it with the threshold to determine whether it meets the standard.
[0057] Part Two:
[0058] Step 1: RGBD Data Acquisition
[0059] An RGBD camera is used to acquire images of the gate tile under inspection, obtaining an RGB image (visible light image) containing color information and a depth image (Depth image) reflecting spatial structure. To ensure the integrity and accuracy of the acquired data, the RGBD camera is kept stable during acquisition to avoid image errors introduced by equipment shaking. A multi-frame strategy is adopted, that is, five consecutive frames of RGB images and their corresponding depth images are acquired from the same viewpoint and position for subsequent inter-frame fusion and noise reduction processing. During acquisition, attention must be paid to controlling the ambient lighting conditions to ensure sufficient and uniform illumination, avoiding images that are too dark or overexposed, which would affect the quality of the visible light image. At the same time, it is also necessary to ensure that the distance between the gate tile and the RGBD camera is within its effective ranging range to ensure the validity and accuracy of the depth image.
[0060] Step 2: RGBD Image Optimization Processing
[0061] The multi-frame RGB images acquired in step 1 are optimized using frame averaging, which means averaging the value of the same pixel across multiple frames to reduce random noise and improve the signal-to-noise ratio.
[0062] The formula for calculating the multi-frame average is as follows:
[0063] Among them: I i(x,y) This represents the pixel value of the i-th frame, where N is the frame number.
[0064] Due to limitations in sensor hardware precision, RGBD cameras may exhibit random fluctuations in depth values at the same pixel location when acquiring depth images, thus affecting the overall measurement accuracy. To improve the stability of depth data and effectively suppress the adverse effects of outliers on subsequent processing steps (such as segmentation and measurement), it is necessary to perform temporal and spatial domain filtering on the depth map.
[0065] Specifically, the five depth images acquired in step 1 can be stacked in chronological order into a 5×W×H three-dimensional matrix, where 5 represents the frame number in the time dimension, and W and H represent the width and height of the image, respectively. Based on this 3D structure, a 5×5×5 three-dimensional median filter operator is applied to each pixel in both the time and spatial dimensions (local neighborhood of the image) for joint filtering. This method can effectively remove random noise and abrupt outliers while preserving edge structure. After filtering, the resulting depth map is smoother and more consistent in spatial structure, and its values are more stable and closer to the true physical depth. This optimization not only improves the overall quality of the depth map but also provides a solid data foundation for the robustness and detection accuracy of subsequent semantic segmentation networks.
[0066] The formula for calculating three-dimensional median filtering is as follows:
[0067] D' (x,y) =median(D (x+i,y+j,t+k) ), i, j, k ∈ {-2, -1, 0, 1, 2}
[0068] Where D (x,y,t) D' represents the depth value at position (x, y) in the t-th frame of the original 3D depth data. (x,y) The filtered depth value is denoted by median(·), which means taking the median value within a 5×5×5 neighborhood. i and j control the spatial neighborhood, and k controls the temporal neighborhood.
[0069] Step 3: Extract the brake shoe segmentation mask
[0070] A semantic segmentation network for gate tiles based on RGBD information is employed, fully integrating the texture features of RGB images with the geometric structure information of depth maps to achieve high-precision segmentation of gate tile regions. RGB images contain rich color, texture, and edge cues, helping the network accurately identify the differences between the gate tiles and the background; while each pixel value in the depth map represents the spatial distance from that point to the camera, effectively reflecting the height variations in the scene, thus strengthening the structural boundary between the gate tiles and the surrounding environment. By fusing these two types of information, the network can maintain stable segmentation results even in scenes with complex backgrounds or significant changes in lighting conditions (see network structure for details). Figure 2 ).
[0071] The semantic segmentation network backbone structure used in this invention is U-net. This network performs stably in image segmentation tasks, possesses excellent multi-scale feature fusion capabilities, and can take into account both global contours and local details, ensuring that the segmentation results have clear edges and accurate shapes. To improve the adaptability of the input data, the RGB images are normalized and standardized before being fed into the network (i.e., the channel mean is subtracted and the standard deviation is divided), while the depth map is normalized to the [0,1] interval, thereby avoiding the influence of different camera ranges and ensuring that the depth features are consistent under different acquisition environments.
[0072] In terms of network architecture design, a dual-channel feature extraction strategy is adopted: RGB images and depth maps are input into two parallel sub-networks for initial feature extraction. Subsequently, the feature maps output by the two are fused along the channel dimension, allowing color and geometric information to fully complement each other. The fused features are then fed into the U-net backbone network for further processing, ultimately generating a gate tile segmentation mask for subsequent analysis. Through this collaborative learning mechanism of RGB and depth information, the network accurately locates the gate tile boundaries while exhibiting stronger robustness and segmentation accuracy, providing a reliable foundation for subsequent thickness calculations.
[0073] Step 4: Wheel center detection
[0074] To obtain a reference point for measuring the brake shoe thickness, the center of the wheel needs to be accurately located in the image. Based on the RGB image preprocessed in step 2, the Hough circle transform algorithm is used for circular target detection to extract the wheel edge and calculate the center coordinates. Considering that the geometric dimensions of the wheel are relatively stable in practical applications, the wheel diameter range and estimated position can be combined as prior information input to effectively constrain the center search range and improve the accuracy and robustness of the center positioning. As the geometric center of the subsequent ray projection, the accuracy of the center positioning directly affects the accuracy of the thickness calculation.
[0075] This invention employs the Hough Circle Transform algorithm to detect wheel contours in denoised RGB images, and then fits the wheel's center and radius using these contours. The basic idea of the Hough Circle Transform is to represent the parameter space of a circle as a three-dimensional space (x, y, ...). c ,y c ,r), where (x c ,y c (x) represents the coordinates of the center of the circle, and r is the radius of the circle. For edge points in the image (x... i ,y i ), satisfying the equation of a circle:
[0076] (x i -x c ) 2 +(y i -y c ) 2 =r 2
[0077] The Hough transform transforms the equation in parameter space by applying the transformation to the equation (x... c ,y c A voting statistics process is performed to select the local maximum value as the parameter for the candidate circle. To improve detection accuracy, the detection process of the wheel center is optimized by incorporating the following prior information: The radius search range of the circle is set [r]. min ,r max The range is preset based on the structural dimensions of the wheel; the approximate center region (x0, y0) is estimated based on the vehicle structure, thereby limiting the search range in the parameter space; multiple candidate circles are evaluated based on edge closure and gradient strength, and only the optimal result is retained as the wheel center. Finally, the wheel center coordinates are (x0, y0). c ,y c This point will serve as the starting point for all subsequent radiation rays.
[0078] Step 5: Project a ray with the center of the wheel as the base point and obtain the line segment of the brake shoe thickness.
[0079] Using the wheel center (x) obtained in step 4 c ,y c Using the origin of the polar coordinate system as θ, and following the angle range of θ∈[0°,359°], 360 rays are generated with a step size of 1 degree. Each ray has a length of l. θ This can be represented as a half-line radii emanating from the center of the circle along direction θ. This ray may intersect the gate segmentation mask M (mask value 255) zero, one, or more times in the image space. To obtain thickness information, the intersection points of each ray with the mask edge need to be precisely calculated: the ray is traversed pixel-level to determine whether the pixels it passes through are located at the mask edge (i.e., undergoing a boundary change with the background 0); if multiple intersection points {P1, P2, ..., P...} exist... n If the first intersection point P is retained, then the first intersection point P is retained. in Intersection with the last point P out The line segment between these two points is the effective thickness range in that direction;
[0080] The pixel length of this line segment can be calculated using Euclidean distance:
[0081]
[0082] The above processing will output a maximum of 360 effective thickness segments, corresponding to 360 directions. To improve the accuracy of thickness measurement, the minimum value (after removing error points) will be selected from these lengths to determine the thinnest part of the brake shoe.
[0083] Step 6: Locate the thinnest part of the brake shoe.
[0084] In step 5, each ray emanating from the center of the wheel... θ For a given thickness segment, all length values with a value of 0 need to be removed, indicating that the ray does not intersect with the mask. The remaining length values are sorted by angle to form the set of thickness distributions in each direction corresponding to the center of the gate shoe:
[0085]
[0086] However, the irregular shapes at both ends of the brake shoe, and the potential for incomplete boundaries or burrs on the mask edges due to segmentation errors, make the thickness calculation at both ends of the brake shoe highly susceptible to errors. Therefore, to improve the accuracy of the thinnest point determination, it is necessary to... The thickness values at both ends are removed.
[0087] From the set Remove certain values near the top and bottom edges of the mask, such as removing the length values in the front and back five directions, i.e., removing {L0,...,L4} and {L n-4 ,...,L n}, to avoid interference caused by incomplete mask boundaries; for the set of remaining thickness values Calculate the minimum value: Record the coordinates P of the intersection point before and after the minimum thickness value. in and P out The line segment marked in this direction is the candidate area for the thinnest part of the brake shoe in the entire structure.
[0088] Step 7: Calculate the brake shoe thickness
[0089] The thinnest pixel thickness L obtained from step 6 min Mapping to the actual physical thickness requires combining depth map information from the RGBD image and calculating the physical spatial distance using the 3D coordinates of the two intersection points on the brake shoe contour line. Let P be the two boundary pixels at the thinnest point. in =(x in ,y in ), P out =(x out ,y out The corresponding depth value is D. in D out Because the RGB and Depth images in an RGBD camera are aligned pixel-by-pixel, the depth map pixels can be reconstructed into 3D point cloud points in the camera coordinate system using the camera intrinsic parameter matrix K. The corresponding intersection point's 3D point cloud coordinates are (X... in ,Y in Z in ), (X out ,Y out Z out The thickness T at the thinnest point of the brake shoe real The calculation formula is as follows:
[0090]
[0091] Finally, T real Safety threshold T for brake shoe thickness thresh The wear is compared, and if the difference is less than a threshold, the brake shoe is considered to have reached its wear limit and should be replaced promptly. This step outputs the final actual physical thickness value used for judgment, and can automatically complete the wear judgment in conjunction with the system's set threshold, providing excellent online detection and alarm capabilities.
[0092] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, for the device embodiments, the above descriptions are merely preferred embodiments of the present invention. Since they are fundamentally similar to the method embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments. The above descriptions are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention, without departing from the principle of the present invention, should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for identifying abnormal thickness of subway train brake shoes based on deep learning, characterized in that... Includes the following steps: Step 1) Acquire images of the target brake shoe to obtain multimodal image data, including RGB images and corresponding depth maps; Step 2) Perform image quality enhancement processing on the multiple frames of RGB images and depth images obtained in Step 1); Step 3) The RGB image and depth image processed in Step 2) are fed into the semantic segmentation network as inputs to perform joint segmentation inference. The gate tile semantic segmentation network based on RGBD information is adopted to fuse the texture features of the RGB image and the geometric structure information of the depth map to achieve high-precision segmentation of the gate tile region. The RGB image contains color, texture and edge cues, and each pixel value in the depth map represents the spatial distance from the point to the camera, reflecting the height undulation in the scene and strengthening the boundary between the gate tile and the surrounding environment at the structural level. Step 4) Use Hough circle transform to find the center of the wheel in the denoised RGB image obtained in step 2; improve the accuracy of center acquisition by using prior information, which includes the wheel diameter range and the center range; Step 5) Based on the center coordinates obtained in Step 4), obtain 360 ray equations with the center as the endpoint at equal intervals. For each ray, find the coordinates of its intersection point with the outline of the gate tile segmentation mask, and calculate the pixel length of each ray within the gate tile segmentation mask. Step 6) Locate the thinnest part of the brake shoe, specifically: 6.1) In step 5), each ray emanating from the center of the wheel... θ For a given thickness segment, remove all length values that are 0, indicating that the ray does not intersect with the mask; 6.2) Sort the remaining length values by angle to form a set of thickness distributions in each direction corresponding to the center of the brake shoe: 6.3) In thickness sets Remove the thickness values at both ends from the set; Remove several values from the mask's edge in the top and bottom extreme directions, such as removing length values in several directions before and after, to obtain the remaining thickness value set. Calculate the minimum value: Record the coordinates P of the intersection point before and after the minimum thickness value. in and P out The line segment marked in this direction is the candidate area for the thinnest part of the brake shoe in the entire structure; Step 7) Calculate the brake shoe thickness and compare it with the threshold to determine whether it meets the standard; 7.1) Combining the depth map information in the RGBD image, the physical spatial distance is calculated using the three-dimensional coordinates of the two intersection points on the brake shoe contour line. The thinnest pixel thickness L obtained in step 6) is then used. min The mapping to the actual physical thickness is as follows: Let P be the two boundary pixels at the thinnest point. in =(x in ,y in ), P out =(x out ,y out The corresponding depth value is D. in D out Using the camera intrinsic parameter matrix K, the pixels in the depth map are reconstructed into 3D point cloud points in the camera coordinate system; the corresponding intersection point point cloud 3D coordinates are (X... in ,Y in Z in ), (X out ,Y out Z out The thickness T at the thinnest point of the brake shoe real The calculation formula is as follows: 7.2) T real Safety threshold T for brake shoe thickness thresh The comparison is performed, and if the value is less than the threshold, it is determined that the brake shoe has reached its wear limit and needs to be replaced in time.
2. The method for identifying abnormal thickness of subway train brake shoes based on deep learning according to claim 1, characterized in that: During the image acquisition process described in step 1), ensure that the camera and the brake shoe surface are facing each other at a positive angle, and continuously acquire multiple frames of image data, that is, continuously acquire multiple frames of RGB images and their corresponding depth images at the same viewpoint and position.
3. The method for identifying abnormal brake shoe thickness in subway trains based on deep learning according to claim 1, characterized in that: Step 2) The image quality enhancement process specifically involves: performing a multi-frame denoising algorithm on the RGB image to suppress random noise; and simultaneously employing temporal and spatial filtering methods on the depth image to remove discrete noise and local defects caused by sensor jitter or acquisition errors.
4. The method for identifying abnormal thickness of subway train brake shoes based on deep learning according to claim 3, characterized in that: The multi-frame denoising algorithm specifically involves averaging the values of the same pixel across multiple frames to reduce random noise and improve the signal-to-noise ratio. The formula for calculating the multi-frame average is as follows: Among them: I i(x,y) This represents the pixel value of the i-th frame, where N is the frame number.
5. The method for identifying abnormal thickness of subway train brake shoes based on deep learning according to claim 3 or 4, characterized in that: The temporal and spatial filtering methods are as follows: Assuming that n frames of depth images were acquired in step 1), they are stacked in chronological order into a three-dimensional matrix with a size of n×W×H, where n represents the number of frames in the time dimension, and W and H represent the width and height of the image, respectively. Based on this three-dimensional structure, a three-dimensional median filter operator of size n×n×n is applied to each pixel in both the time and spatial dimensions to perform joint filtering. The formula for calculating three-dimensional median filtering is as follows: D' (x,y) =median(D (x+i,y+j,t+k) ),i,j,k∈{-2,-1,0,1,2} Where D (x,y,t) D' represents the depth value at position (x, y) in the t-th frame of the original 3D depth data. (x,y) The filtered depth value is denoted by median(·), which means taking the median value within an n×n×n neighborhood. i and j control the spatial neighborhood, and k controls the temporal neighborhood.
6. The method for identifying abnormal thickness of subway train brake shoes based on deep learning according to claim 1, characterized in that: Step 3) describes the process of obtaining the gate shoe segmentation mask through semantic segmentation using the U-net network backbone structure: 3.1) Before being fed into the network, the RGB images are normalized and standardized, while the depth map is normalized to the [0,1] interval to ensure that the depth features are consistent under different acquisition environments; 3.2) Design the network structure and adopt a dual-channel feature extraction strategy: RGB image and depth map are respectively input into two parallel sub-networks for preliminary feature extraction; 3.3) The feature maps output by the two are fused along the channel dimension to make the color and geometric information fully complementary; 3.4) The features fused in step 3.3) are fed into the U-net backbone network for further processing, and finally a gate shoe segmentation mask is generated for subsequent analysis.
7. The method for identifying abnormal thickness of subway train brake shoes based on deep learning according to claim 1, characterized in that: Step 4) The Hough circle transformation algorithm specifically involves: representing the parameter space of the circle as a three-dimensional space (x... c ,y c ,r), where (x c ,y c (x) represents the coordinates of the center of the circle, and r is the radius of the circle; for edge points in the image (x) i ,y i ), satisfying the equation of a circle: (x i -x c ) 2 +(y i -y c ) 2 =r 2 For (x) that satisfy this equation in the parameter space c ,y c We perform a voting statistics on r) and select the local maximum value as the parameter of the candidate circle.
8. The method for identifying abnormal brake shoe thickness in subway trains based on deep learning according to claim 1 or 7, characterized in that: Step 4) The method for improving the accuracy of center acquisition specifically involves: setting the radius search range of the circle [r] min ,r max The range is preset based on the structural dimensions of the wheel; the approximate center region (x0, y0) is estimated based on the vehicle structure, thereby limiting the search range of the parameter space; multiple candidate circles are evaluated based on edge closure and gradient strength, and only the optimal result is retained as the wheel center; finally, the wheel center coordinates are obtained as (x0, y0). c ,y c This point will serve as the starting point for all subsequent radiation rays.
9. The method for identifying abnormal thickness of subway train brake shoes based on deep learning according to claim 1, characterized in that: Step 5) The process of calculating the pixel length of each ray within the gate tile segmentation mask is specifically as follows: Each ray l θ It is represented as a half-line starting from the center of the circle along the direction θ. This ray intersects the gate segmentation mask M 0, 1 or more times in the image space. The ray is traversed at the pixel level to determine whether the pixel it passes through is located at the edge of the mask. If there exist multiple intersection points {P1, P2, ..., P...} n If the first intersection point P is retained, then the first intersection point P is retained. in Intersection with the last point P out The line segment between these two points represents the effective thickness range in that direction; the pixel length L of this line segment... θ Calculated using Euclidean distance: