Real-time Track Crack Detection Method and System Based on Multidimensional Machine Vision
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-07
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]为解决上述技术问题,提供一种基于多维机器视觉的轨道裂纹实时检测方法及系统,本技术方案解决了上述背景技术中提出的难以实现像素级的精准对齐和难以有效适应轨道现场复杂多变的光照条件与背景干扰,导致裂纹检测的鲁棒性不足问题
本发明提供一种基于多维机器视觉的轨道裂纹实时检测方法及系统,通过编码器同步触发多视角工业相机与线激光轮廓传感器,实现二维图像与三维深度数据的像素级时空对齐,构建融合二维纹理特征与三维深度特征的裂纹多维特征图,结合语义分割网络精准提取轨道表面有效区域,利用全卷积裂纹检测网络实现裂纹的精准识别与候选区域提取,并通过Zhang-Suen细化算法与切线法线方向宽度测量实现裂纹长度、宽度、深度的自动化精准量化,最终依据裂纹几何参数完成三级等级判定并与里程信息关联生成可追溯检测记录,形成从数据采集、特征融合、裂纹识别到参数量化的全流程自动化闭环检测,有效克服了单一视觉信息局限性、多源数据时空不同步、光照背景干扰严重、裂纹量化能力不足及检测流程碎片化等现有技术缺陷,有效提升了轨道裂纹检测的准确率、鲁棒性与运维效率;进一步地,本方案通过裂纹发展时期判别与自适应特征融合权重调整,解决了不同阶段裂纹特征敏感性差异的技术问题;通过非对称卷积核组合增强了对细长裂纹结构的感知能力;通过形态约束损失函数使网络输出更符合裂纹的连续平滑物理形态,有效提升了裂纹检测的准确性与模型泛化能力。
Smart Images

Figure CN122573808A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of track crack detection, specifically to a real-time track crack detection method and system based on multi-dimensional machine vision. Background Technology
[0002] As a crucial component of modern urban and regional transportation, rail transit's operational safety directly impacts the safety of people's lives and property, as well as the stable operation of the social economy. The track, as the core load-bearing structure of the rail transit system, endures the combined effects of multiple factors, including train dynamic loads, temperature stress, environmental corrosion, and material fatigue. Fatigue cracks are prone to develop on the rail top, rail sides, and weld areas. If these cracks are not detected and effectively addressed in time, they may gradually propagate with the continuous loading of the train, eventually leading to rail breaks and serious accidents such as derailment and overturning. Therefore, efficient and accurate detection and monitoring of track cracks is a critical aspect of ensuring the safe operation of rail transit.
[0003] The core detection principle of the monorail inspection robot lies in the integration of multi-dimensional perception technology to achieve automated identification of track defects. The robot controls an industrial camera and a laser contour sensor mounted on its body to collect data synchronously. The industrial camera is responsible for acquiring high-resolution two-dimensional images of the rail top, rail side, and weld seam areas to identify the linear texture features of cracks. The line laser contour sensor, on the other hand, emits line structured light and receives reflected signals to obtain continuous three-dimensional contour data of the track cross-section, which is used to extract the depth features of cracks. By aligning the two-dimensional images and three-dimensional depth data in time and space and fusing features, the robot can comprehensively utilize the dual abnormal information of cracks in image texture and surface morphology to effectively distinguish cracks from visually similar interference objects such as rust, scratches, and oil stains. Furthermore, it can calculate the geometric parameters such as the length, width, and depth of cracks, providing accurate quantitative basis for track maintenance decisions.
[0004] However, in existing technologies, although monorail inspection robots are generally equipped with visual acquisition devices such as industrial cameras and laser contour sensors, the two-dimensional image acquisition and three-dimensional laser scanning in current detection schemes often operate independently, lacking a unified synchronous triggering mechanism. This leads to deviations in spatial position and temporal sequence between image data and depth data, making it difficult to achieve pixel-level precise alignment, which directly affects the reliability of multi-dimensional feature fusion and the accuracy of crack identification. Furthermore, existing image processing technologies often employ fixed thresholds or simple filtering methods, which are difficult to effectively adapt to the complex and variable lighting conditions and background interference at the track site, resulting in insufficient robustness in crack detection. Summary of the Invention
[0005] To address the aforementioned technical problems, a real-time track crack detection method and system based on multi-dimensional machine vision is provided. This technical solution solves the problem mentioned in the background technology that it is difficult to achieve pixel-level precise alignment and difficult to effectively adapt to the complex and ever-changing lighting conditions and background interference at the track site, resulting in insufficient robustness of crack detection.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A real-time track crack detection method based on multidimensional machine vision includes: The robot's walking displacement data is collected by the encoder, and a synchronous trigger signal is generated according to the preset displacement interval, which simultaneously triggers the data acquisition of the industrial camera and the line laser contour sensor. The acquired images of the rail top, rail side and weld area are spatiotemporally aligned with the rail section depth data to form a fused data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix. Illumination compensation and noise suppression are performed on the two-dimensional unfolded image to obtain a standardized image, which is then input into a semantic segmentation network to extract the effective region of the track surface. Within the effective area of the track surface, the gray-level gradient distribution features of the two-dimensional unfolded image and the local height change features of the three-dimensional depth data are extracted respectively. The two types of features are fused at the pixel level to construct a multi-dimensional feature map of the crack. The multidimensional feature map of the crack is input into the crack detection network, which outputs the probability distribution map of the existence of cracks, and extracts the pixel-level location information of the crack candidate region based on the probability threshold. Based on the pixel-level location information of the crack candidate region, the actual length, maximum width and average depth of the crack are calculated and associated with the current mileage information to generate a detection record containing crack parameters and location.
[0007] Preferably, the step of acquiring robot walking displacement data through an encoder and generating a synchronous trigger signal according to a preset displacement interval, simultaneously triggering data acquisition from the industrial camera and the line laser contour sensor, specifically includes: Control the monorail inspection robot to move at a constant speed along the monorail track to be tested; The quadrature pulse signal output by the encoder is acquired and converted into a real-time walking displacement value by a counter; When the real-time walking displacement value reaches the preset displacement interval, a trigger pulse is generated and simultaneously output to the industrial cameras and line laser contour sensors arranged in top view, side view, and oblique view. After receiving a trigger pulse, the industrial camera and the line laser profile sensor simultaneously perform a single image acquisition and cross-sectional profile scanning. The real-time walking displacement value corresponding to the trigger time is used as the mileage identifier of the data frame collected this time, and is bound and stored with the collected multi-view images and cross-sectional contour data.
[0008] Preferably, the step of spatiotemporally aligning the acquired images of the rail top, rail side, and weld area with the rail cross-section depth data to form a fused data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix specifically includes: The images of the rail top, rail side and weld area at the same trigger moment are registered according to the preset spatial splicing template to generate a two-dimensional unfolded view of the entire rail section. Extract the cross-sectional contour point cloud output by the line laser contour sensor, and generate a depth matrix corresponding to the pixel coordinates of the two-dimensional unfolded image through a linear interpolation algorithm; Using mileage identifiers as the primary key, the two-dimensional unfolded image and depth matrix are encapsulated into a fused data frame to achieve point-to-point correspondence between image pixels and depth values.
[0009] Preferably, the step of performing illumination compensation and noise suppression processing on the two-dimensional unfolded image to obtain a standardized image, and then inputting the standardized image into a semantic segmentation network to extract the effective region of the track surface specifically includes: A two-dimensional unfolded image is obtained from the fused data frame. Adaptive histogram equalization is used to adjust the local contrast of the image. Bilateral filtering is used to retain edge information while suppressing high-frequency noise to obtain a standardized image. The processed, normalized image is input into a lightweight semantic segmentation network with an encoder-decoder structure, and the output is a binary mask of the track surface region; A binary mask is used to filter pixels in the standardized image, retaining pixels in the orbital region to obtain an image of the effective region on the orbital surface.
[0010] Preferably, the step of extracting the gray-level gradient distribution features of the two-dimensional unfolded image and the local height abrupt change features of the three-dimensional depth data within the effective area of the track surface, and fusing the two types of features at the pixel level to construct a multi-dimensional crack feature map specifically includes: Within the effective area of the track surface, a multi-directional Gabor filter bank is used to extract texture responses at different angles to generate a two-dimensional texture feature map; The depth matrix is subjected to first-order difference operation to extract the height change rate along the track direction and perpendicular to the track direction, and the local depth anomaly is calculated. After integration, a three-dimensional morphological feature map is generated. The two-dimensional texture feature map and the three-dimensional morphological feature map are weighted and superimposed according to pixel position to construct a multi-dimensional feature map of cracks.
[0011] Preferably, the step of inputting the multidimensional feature map of the crack into the crack detection network, outputting the probability distribution map of the existence of the crack, and extracting the pixel-level location information of the crack candidate region based on the probability threshold specifically includes: The crack detection network adopts a fully convolutional structure. It takes the multi-dimensional feature map of cracks as input, extracts contextual semantic information through multi-scale convolution, and after passing through the Sigmoid activation function, outputs a probability map with the same resolution as the input image, as well as the crack recognition confidence of the two-dimensional texture feature map and the three-dimensional morphological feature map. The multi-scale convolution includes a combination of asymmetric convolution kernels, and the value at each pixel position in the probability map represents the probability value of that pixel belonging to the crack category. The crack development period is divided into the initiation stage, the propagation stage, and the penetration stage. Based on the crack recognition confidence of the two feature maps, the fusion weight of the two feature maps is dynamically adjusted. The training loss function and morphological constraint evaluation index of the crack detection network are constructed. The training loss function is a weighted sum of cross-entropy loss and Dice loss. The morphological constraint evaluation index is a weighted sum of connectivity loss and curvature smoothing loss. The probability map is binarized according to a preset probability threshold. Pixels with probability values greater than or equal to the threshold are marked as crack candidate pixels, and pixels with probability values less than the threshold are marked as background pixels, thus obtaining the initial crack segmentation mask. Connectivity analysis is performed on the initial crack segmentation mask to remove isolated noise points with an area smaller than a preset area threshold, while retaining the crack candidate region and its pixel-level position coordinates.
[0012] Preferably, the step of calculating the actual length, maximum width, and average depth of the crack based on the pixel-level location information of the crack candidate region, and associating it with the current mileage information to generate a detection record containing crack parameters and location specifically includes: Obtain the crack candidate regions and their pixel-level position coordinates after connected component analysis, where pixel value 1 corresponds to the crack candidate region and pixel value 0 corresponds to the background region. The Zhang-Suen thinning algorithm is used to iteratively thin the binary image of the crack candidate region. It traverses all crack pixels with a pixel value of 1 in the image and determines whether each pixel meets the preset thinning conditions. Pixels that meet the thinning conditions are marked and deleted, while pixels that do not meet the conditions are kept. The traversal and deletion steps are repeated until no pixels are marked or deleted in the current traversal, at which point the iteration terminates. After the iteration terminates, the remaining pixels are connected sequentially according to their pixel-level position coordinates to form a single-pixel-width line, which is the crack center line. At the same time, all pixel-level position coordinates of the crack center line are recorded. Traverse all adjacent pixels along the crack centerline, calculate the Euclidean distance between each pair of adjacent pixels, sum them up to obtain the centerline pixel length, and multiply by the image resolution to obtain the actual length of the crack. For each pixel on the crack centerline, calculate the tangent direction of the centerline at that point, and the normal direction is the direction perpendicular to the tangent direction. Scan along the normal direction to both sides, count the number of consecutive crack pixels, multiply by the image resolution to obtain the crack width at that point, and take the maximum width among all centerline points as the maximum crack width. Obtain the depth value at the corresponding position in the depth matrix, and calculate the arithmetic mean of the depth values of all pixels in the crack candidate region as the average crack depth. The length, width, and depth parameters of the crack are bound to the mileage identifier of the current fused data frame, and a detection record is generated and uploaded to the monitoring platform according to the preset crack level classification rules.
[0013] Furthermore, this solution proposes a real-time track crack detection system based on multi-dimensional machine vision to implement the aforementioned real-time track crack detection method based on multi-dimensional machine vision, including: The data acquisition module is used to acquire robot walking displacement data through an encoder and generate a synchronous trigger signal according to a preset displacement interval, which simultaneously triggers data acquisition from the industrial camera and the line laser contour sensor. The data preprocessing module is used to spatiotemporally align the acquired images of the rail top, rail side, and weld area with the track cross-section depth data to form a fused data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix; perform illumination compensation and noise suppression processing on the two-dimensional unfolded image to obtain a standardized image, and input the standardized image into a semantic segmentation network to extract the effective area of the track surface; within the effective area of the track surface, extract the gray-level gradient distribution features of the two-dimensional unfolded image and the local height abrupt change features of the three-dimensional depth data respectively, and perform pixel-level fusion of the two types of features to construct a multi-dimensional feature map of the crack; The crack recognition module is used to input the multi-dimensional feature map of the crack into the crack detection network, output the probability distribution map of the existence of the crack, and extract the pixel-level position information of the crack candidate region based on the probability threshold; based on the pixel-level position information of the crack candidate region, calculate the actual length, maximum width and average depth of the crack, and associate them with the current mileage information to generate a detection record containing crack parameters and position.
[0014] Preferably, the data preprocessing module includes: The fusion data unit is used to align the acquired images of the rail top, rail side and weld area with the rail section depth data in time and space to form a fusion data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix. The effective region unit is used to perform illumination compensation and noise suppression processing on the two-dimensional unfolded image to obtain a standardized image, and input the standardized image into the semantic segmentation network to extract the effective region of the track surface. The feature map unit is used to extract the gray-level gradient distribution features of the two-dimensional unfolded map and the local height change features of the three-dimensional depth data within the effective area of the track surface, respectively, and to fuse the two types of features at the pixel level to construct a multi-dimensional feature map of the crack.
[0015] Preferably, the crack identification module includes: A crack identification unit is used to input a multi-dimensional feature map of cracks into a crack detection network, output a probability distribution map of crack existence, and extract pixel-level location information of crack candidate regions based on a probability threshold. A crack recording unit is used to calculate the actual length, maximum width, and average depth of a crack based on the pixel-level location information of a crack candidate region, and associate it with the current mileage information to generate a detection record containing crack parameters and location.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a real-time track crack detection method and system based on multi-dimensional machine vision. It achieves pixel-level spatiotemporal alignment of two-dimensional images and three-dimensional depth data by synchronously triggering a multi-view industrial camera and a line laser contour sensor through an encoder. This constructs a multi-dimensional crack feature map integrating two-dimensional texture features and three-dimensional depth features. A semantic segmentation network is used to accurately extract the effective area of the track surface. A fully convolutional crack detection network is used to accurately identify cracks and extract candidate regions. The Zhang-Suen thinning algorithm and tangent normal direction width measurement are used to automatically and accurately quantify the crack length, width, and depth. Finally, based on the crack geometric parameters, a three-level crack level determination is completed and associated with mileage information to generate a traceable detection record, forming a data... This fully automated closed-loop detection system, encompassing data acquisition, feature fusion, crack identification, and parameter quantization, effectively overcomes existing technological limitations such as the constraints of single visual information, spatiotemporal asynchrony of multi-source data, severe interference from lighting and background, insufficient crack quantization capabilities, and fragmented detection processes. This significantly improves the accuracy, robustness, and operational efficiency of track crack detection. Furthermore, this solution addresses the technical challenge of varying crack feature sensitivity at different stages by using crack development stage discrimination and adaptive feature fusion weight adjustment. The system enhances its ability to perceive slender crack structures through asymmetric convolutional kernel combinations. Finally, the morphological constraint loss function makes the network output more consistent with the continuous and smooth physical morphology of cracks, effectively improving the accuracy of crack detection and the model's generalization ability. Attached Figure Description
[0017] Figure 1This is a flowchart of the real-time track crack detection method based on multi-dimensional machine vision of the present invention. Figure 2 This is a flowchart illustrating the process of forming a fused data frame comprising a two-dimensional unfolded map and a three-dimensional depth matrix according to the present invention. Figure 3 This is a flowchart illustrating the extraction of the effective area of the track surface according to the present invention. Figure 4 This is a flowchart of the present invention, which inputs a multidimensional feature map of cracks into a crack detection network and outputs a probability distribution map of crack existence. Figure 5 The flowchart for generating detection records containing crack parameters and locations is provided for this invention. Detailed Implementation
[0018] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0019] Reference Figure 1 As shown, a real-time track crack detection method based on multidimensional machine vision includes: The robot's walking displacement data is collected by the encoder, and a synchronous trigger signal is generated according to the preset displacement interval, which simultaneously triggers the data acquisition of the industrial camera and the line laser contour sensor. The acquired images of the rail top, rail side and weld area are spatiotemporally aligned with the rail section depth data to form a fused data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix. Illumination compensation and noise suppression are performed on the two-dimensional unfolded image to obtain a standardized image, which is then input into a semantic segmentation network to extract the effective region of the track surface. Within the effective area of the track surface, the gray-level gradient distribution features of the two-dimensional unfolded image and the local height change features of the three-dimensional depth data are extracted respectively. The two types of features are fused at the pixel level to construct a multi-dimensional feature map of the crack. The multidimensional feature map of the crack is input into the crack detection network, which outputs the probability distribution map of the existence of cracks, and extracts the pixel-level location information of the crack candidate region based on the probability threshold. Based on the pixel-level location information of the crack candidate region, the actual length, maximum width and average depth of the crack are calculated and associated with the current mileage information to generate a detection record containing crack parameters and location.
[0020] This solution utilizes encoder displacement interval triggering to enable multi-view industrial cameras and line laser contour sensors to synchronously perform image acquisition and cross-sectional scanning at the same spatial position on the track. The displacement value at the trigger moment is used as a mileage marker and bound to the acquired data for storage, ensuring that a unified mileage benchmark is established between the 2D image and 3D depth data during the acquisition phase, avoiding spatiotemporal errors introduced by post-alignment. The multi-view images are stitched together to form a 2D unfolded map of the entire track cross-section, which is then encapsulated with a depth matrix generated through linear interpolation into a fused data frame, providing pixel-level aligned texture information for subsequent crack detection. A lightweight semantic segmentation network extracts the effective area of the track surface, effectively eliminating background interference from the track bed, fasteners, etc., reducing the interference of non-track features on crack identification. Simultaneously, within the effective area of the track, [the solution extracts...]. Multi-directional Gabor texture features and local depth anomaly features are weighted and superimposed according to pixel position to construct a multi-dimensional feature map of cracks. The linear texture anomalies of cracks in the image and the indentation anomalies in the depth form complementary discrimination criteria, which improves the ability to distinguish cracks from interference such as rust and scratches. The crack probability distribution map is output through a fully convolutional crack detection network. Crack candidate regions are extracted by probability threshold binarization and connected component analysis. The crack center line is extracted by Zhang-Suen thinning algorithm. Combined with the tangent normal direction, the crack length and width are accurately measured. The average crack depth is obtained from the depth matrix. The crack geometric parameters are bound with the mileage marker to generate a traceable detection record, realizing a closed-loop detection process from data acquisition, feature fusion, crack recognition to parameter quantization.
[0021] The process of acquiring robot displacement data via encoder and generating synchronous trigger signals based on preset displacement intervals to simultaneously trigger data acquisition from industrial cameras and line laser contour sensors specifically includes: Control the monorail inspection robot to move at a constant speed along the monorail track to be tested; The quadrature pulse signal output by the encoder is acquired and converted into a real-time walking displacement value by a counter; When the real-time walking displacement value reaches the preset displacement interval, a trigger pulse is generated and simultaneously output to the industrial cameras and line laser contour sensors arranged in top view, side view, and oblique view. After receiving a trigger pulse, the industrial camera and the line laser profile sensor simultaneously perform a single image acquisition and cross-sectional profile scanning. The real-time walking displacement value corresponding to the trigger time is used as the mileage identifier of the data frame collected this time, and is bound and stored with the collected multi-view images and cross-sectional contour data.
[0022] The encoder is installed on the end of the walking wheel axle and outputs orthogonal pulses as the robot moves along the track. The pulse counter accumulates the number of pulses in real time and converts them into displacement values based on the wheel diameter. When the accumulated displacement reaches a preset interval, the controller generates an edge trigger signal, which simultaneously drives all cameras and laser profilometers to collect track data. The real-time walking displacement value corresponding to the trigger moment is used as the mileage identifier of the data frame collected this time. This data is bound and stored with the collected multi-view images and cross-sectional contour data to ensure that each sensor completes data collection at the same spatial position on the track, laying the foundation for subsequent spatiotemporal alignment of data. The preset displacement interval is determined based on the field of view width of the industrial camera along the track walking direction. The preset displacement interval is less than or equal to the field of view width to ensure that there is a preset overlap rate between adjacent collected images. The overlap rate is set to 10% to 30% to ensure the continuity of detection coverage while avoiding data redundancy.
[0023] Reference Figure 2 As shown, the process of forming a fused data frame containing a two-dimensional unfolded map and a three-dimensional depth matrix specifically includes: The images of the rail top, rail side and weld area at the same trigger moment are registered according to the preset spatial splicing template to generate a two-dimensional unfolded view of the entire rail section. Extract the cross-sectional contour point cloud output by the line laser contour sensor, and generate a depth matrix corresponding to the pixel coordinates of the two-dimensional unfolded image through a linear interpolation algorithm; Using mileage identifiers as the primary key, the two-dimensional unfolded image and depth matrix are encapsulated into a fused data frame to achieve point-to-point correspondence between image pixels and depth values.
[0024] This can be explained by the fact that cameras with different perspectives have fixed geometric relationships in their physical installation positions. Multiple images can be mapped to a unified track unfolding coordinate system through a pre-calibrated stitching template. The cross-sectional point cloud output by the line laser profile sensor contains the relative height information of each position in the lateral direction of the track. The discrete point cloud is mapped onto a grid with the same resolution as the two-dimensional unfolded map through a linear interpolation algorithm to form a depth matrix. Using the mileage marker as a unique index, the two-dimensional unfolded map is bound to the depth matrix to ensure that each pixel has a corresponding depth value, thus achieving precise spatiotemporal alignment between the two-dimensional unfolded map and the three-dimensional depth data. The preset spatial splicing template is pre-calibrated and set through the following steps: Obtain the intrinsic parameter matrices and distortion coefficients of the top-view, side-view, and oblique-view industrial cameras respectively; The intrinsic parameter matrix and distortion coefficients were calculated using the Zhang Zhengyou calibration method after being photographed from multiple angles within the field of view of each camera through a checkerboard calibration plate. Establish a global unfolded coordinate system for the monorail track, with the track centerline as the X-axis, the transverse direction of the track as the Y-axis, and the vertical direction upward from the top plane of the track as the Z-axis. Set the imaging resolution under the global unfolded coordinate system according to the crack detection accuracy requirements. Determine the corresponding position intervals of the rail top, rail side and weld area in the global unfolded coordinate system. The rail top area corresponds to the central interval of the Y-axis, the rail side area corresponds to the intervals on both sides of the Y-axis, and the weld area is located in the connection transition area between the rail top and the rail side. The calibration plate is placed at multiple preset feature positions on the track surface (covering the track top, track side and weld area). The external parameters of each camera are calibrated by the calibration plate, and the rotation matrix and translation vector of each camera's imaging plane to the global unfolded coordinate system are obtained. Based on the rotation matrix, translation vector and preset imaging resolution, a mapping function is established from the pixel coordinates of the images acquired by each camera to the coordinates of the global unfolded coordinate system. The mapping function is stored in the form of a lookup table to form a spatial stitching template. Spatial stitching templates are used to map local track images acquired by cameras from different perspectives and positions to the same global unfolded coordinate system, eliminating spatial offsets caused by differences in camera installation positions. This allows images of the rail top, rail side, and weld seam areas to be stitched together into a complete two-dimensional unfolded view of the track cross section, providing a unified coordinate reference for the point-to-point correspondence between the subsequent two-dimensional unfolded view and three-dimensional depth data.
[0025] Reference Figure 3 As shown, the effective area extracted from the track surface specifically includes: A two-dimensional unfolded image is obtained from the fused data frame. Adaptive histogram equalization is used to adjust the local contrast of the image. Bilateral filtering is used to retain edge information while suppressing high-frequency noise to obtain a standardized image. The processed, normalized image is input into a lightweight semantic segmentation network with an encoder-decoder structure, and the output is a binary mask of the track surface region; A binary mask is used to filter pixels in the standardized image, retaining pixels in the orbital region to obtain an image of the effective region on the orbital surface.
[0026] This can be explained by the fact that the lighting conditions at the track site change drastically. Adaptive histogram equalization can independently stretch the contrast according to the brightness distribution of the local area of the image, improving the distinction between cracks and background. Bilateral filtering smooths noise while keeping the edges clear, avoiding blurring of crack features. The semantic segmentation network adopts a lightweight structure to reduce computational overhead while ensuring accuracy. After outputting the track area mask, irrelevant background is removed through masking operations, focusing on the track surface. The lightweight semantic segmentation network of the encoder-decoder structure outputs a binary mask for the track surface region, specifically comprising: A lightweight semantic segmentation network is constructed using an encoder-decoder structure based on MobileNetV3 as the backbone network, with preprocessed track-normalized images as network input. The encoder part extracts multi-scale features from the input image through depthwise separable convolution, and outputs low, medium and high-dimensional feature maps in sequence. At the same time, it strengthens the features of the track region and suppresses the background features through an attention mechanism. The decoder uses transposed convolution to upsample the high-dimensional feature map, gradually restoring the image resolution, and fuses the multi-scale feature map output by the encoder to make up for the feature loss during the upsampling process. The network output layer uses the Sigmoid activation function to perform binary classification on each pixel in the image and outputs a binary mask with the same resolution as the input image. The area with a pixel value of 1 corresponds to the effective area of the orbital surface, and the area with a pixel value of 0 corresponds to the invalid area of the background. During the network training phase, labeled images under multiple track scenarios (different lighting, different wear levels, different contaminants) are used as the training set. The cross-entropy loss function is used as the optimization objective. The network is iteratively trained using the gradient descent algorithm until the network segmentation accuracy reaches a preset threshold (not less than 98%). After training, the network is pre-deployed to the computing unit of the monorail robot and outputs a binary mask of the track surface region in real time. The lightweight semantic segmentation network reduces the number of model parameters and computational cost through the MobileNetV3 backbone network, adapting to the real-time processing requirements of the edge computing unit of the monorail robot; the encoder is responsible for feature extraction, and the decoder is responsible for resolution restoration and feature fusion, ensuring accurate segmentation of the track surface area (track top, track side, weld seam); the binary mask can be directly used for subsequent ROI region extraction, removing background interference such as track bed and debris, providing a clean image area for crack feature extraction; The step of using a binary mask to filter pixels in a standardized image, retaining pixels in the orbital region, and obtaining an effective region image of the orbital surface specifically includes: The preprocessed standardized image is mapped pixel-by-pixel to a binary mask of the orbital region at the same resolution; Iterate through each pixel position in the image. If the pixel value at that position is a valid identifier in the binary mask, retain the pixel value of the original image at that position. If the pixel value at that position is an invalid identifier in the binary mask, set the pixel value of the original image at that position to zero. After pixel-by-pixel determination and assignment, an effective area image of the track surface is obtained, which retains only the track surface area and the rest of the area is a black background; In the binary mask, the area with a pixel value of 1 is marked as the effective area of the track surface, and the area with a pixel value of 0 is marked as the invalid background area. By performing pixel-by-pixel AND operation to match and filter the original image with the mask, invalid pixels such as track bed, surrounding environment, and interfering objects can be accurately removed, and only the effective track area where the track top, track side and weld are located can be retained. This reduces the amount of data and interference factors in subsequent multi-dimensional feature extraction, and improves the efficiency and accuracy of crack detection.
[0027] The step of extracting the gray-level gradient distribution features of the two-dimensional unfolded image and the local height abrupt change features of the three-dimensional depth data within the effective area of the track surface, and fusing the two types of features at the pixel level to construct a multi-dimensional crack feature map specifically includes: Within the effective area of the track surface, a multi-directional Gabor filter bank is used to extract texture responses at different angles to generate a two-dimensional texture feature map; The depth matrix is subjected to first-order difference operation to extract the height change rate along the track direction and perpendicular to the track direction, and the local depth anomaly is calculated. After integration, a three-dimensional morphological feature map is generated. The two-dimensional texture feature map and the three-dimensional morphological feature map are weighted and superimposed according to pixel position to construct a multi-dimensional feature map of cracks.
[0028] This can be explained by the fact that cracks appear as linear textures at specific angles in an image. Gabor filters exhibit strong responses in the corresponding angular directions. Simultaneously, the crack region exhibits local depth depressions, and depth difference operations can extract locations of abrupt height changes. Weighted superposition of these two types of features significantly enhances regions that simultaneously satisfy both texture linearity and depth depression anomalies in the feature map, thereby improving the accuracy and anti-interference capability of crack detection. The multi-directional Gabor filter bank extracts texture responses at different angles, specifically including: Construct a Gabor filter bank with 8 preset directions, with an angular spacing of 22.5 degrees between each filter, covering the texture response range from 0 degrees to 157.5 degrees respectively. The filter wavelength is set to 4 pixels and the standard deviation of the Gaussian envelope is set to 1.5. The effective area image of the track surface is input into the Gabor filter bank, and convolution operation is performed sequentially through the Gabor filter in each direction to generate the texture response map of the corresponding angle. Amplitude normalization is performed on the texture response maps from various angles to enhance the texture contrast in high-response areas, thereby obtaining multi-dimensional directional texture features and forming a two-dimensional texture feature map. Gabor filter banks simulate the receptive field characteristics of the human visual system, exhibiting specific response capabilities to linear and striped textures in different directions. Multi-directional filter banks can simultaneously capture the texture variation features of track surface cracks in the transverse, longitudinal, and oblique directions, converting a single image into a feature set with multi-directional responses, providing key texture information for subsequent crack feature and environmental interference features. The calculation of local depth anomaly specifically includes: Centered on each pixel in the depth matrix, select a local neighborhood window of a preset size and traverse all valid track region pixels in the depth matrix; Calculate the mean depth and standard deviation of depth of all pixels within the local neighborhood window, and use the mean depth as the reference depth of the current local region. The difference between the current center pixel's depth value and the reference depth is calculated to obtain the absolute value of the depth deviation. When the depth standard deviation is greater than zero, the ratio of the absolute value of the depth deviation to the depth standard deviation is used as the local depth anomaly of the pixel; when the depth standard deviation is equal to zero, the local depth anomaly of the pixel is zero. After traversal, the local depth anomalies of each pixel are arranged according to pixel position to generate a 3D morphological feature map with the same resolution as the depth matrix. Local depth anomaly is used to quantify the degree of abrupt change in the unevenness of a point on the track surface relative to the surrounding area. Defective areas such as cracks, spalling, and depressions will show a significantly higher local depth anomaly than normal areas. Local depth anomaly can effectively distinguish normal track surfaces from defective areas.
[0029] Reference Figure 4 As shown, the step of inputting the multidimensional feature map of the crack into the crack detection network and outputting the probability distribution map of the existence of the crack specifically includes: The crack detection network adopts a fully convolutional structure. It takes the multi-dimensional feature map of cracks as input, extracts contextual semantic information through multi-scale convolution, and after passing through the Sigmoid activation function, outputs a probability map with the same resolution as the input image, as well as the crack recognition confidence of the two-dimensional texture feature map and the three-dimensional morphological feature map. The multi-scale convolution includes a combination of asymmetric convolution kernels, and the value at each pixel position in the probability map represents the probability value of that pixel belonging to the crack category. The crack development period is divided into the initiation stage, the propagation stage, and the penetration stage. Based on the crack recognition confidence of the two feature maps, the fusion weight of the two feature maps is dynamically adjusted. The training loss function and morphological constraint evaluation index of the crack detection network are constructed. The training loss function is a weighted sum of cross-entropy loss and Dice loss. The morphological constraint evaluation index is a weighted sum of connectivity loss and curvature smoothing loss. The probability map is binarized according to a preset probability threshold. Pixels with probability values greater than or equal to the threshold are marked as crack candidate pixels, and pixels with probability values less than the threshold are marked as background pixels, thus obtaining the initial crack segmentation mask. Connectivity analysis is performed on the initial crack segmentation mask to remove isolated noise points with an area smaller than a preset area threshold, while retaining the crack candidate region and its pixel-level position coordinates.
[0030] This can be explained by the fact that fully convolutional networks can accept inputs of any size and output pixel-level classification results, multi-scale convolution can simultaneously capture the slender structure of cracks and surrounding contextual information, probability thresholds can be adjusted according to the precision and recall requirements of the actual scene, and connected component analysis is used to filter out small false detection areas caused by noise and retain real crack candidates. The crack detection network adopts a fully convolutional structure, taking the multi-dimensional feature map of the crack as input, and extracts contextual semantic information through multi-scale convolution, specifically including: A fully convolutional crack detection network is constructed, which consists of a feature extraction module, a multi-scale fusion module, and an output module, with the multi-dimensional feature map of the crack as the only input to the network. The feature extraction module uses a combination of asymmetric convolution kernels. It uses 1×5 and 5×1 asymmetric convolution kernels to extract the fine texture features in the horizontal and vertical directions, and uses 1×7 and 7×1 asymmetric convolution kernels to extract the continuous direction features of long cracks. It performs multi-scale convolution operations on the multi-dimensional feature map of cracks to extract shallow detail features, mid-level texture features and deep semantic features in sequence. The multi-scale fusion module performs channel concatenation on the multi-scale feature maps output by each asymmetric convolution kernel, and performs channel compression through a 1×1 convolution kernel to fuse contextual semantic information at different scales, suppress invalid interference features, and enhance crack features. After multi-scale convolution and feature fusion, the output module outputs a fused feature map containing crack context semantic information, providing accurate feature support for subsequent crack region localization and identification. During the network training phase, multi-dimensional feature maps containing different types and sizes of cracks and various interference scenarios are used as the training set. The weighted sum of cross-entropy loss function and Dice loss function is used as the optimization objective, and the Adam optimization algorithm is used for iterative training. The fully convolutional structure eliminates the need for image size compression and restoration, allowing direct adaptation to multi-dimensional feature maps of cracks at any resolution and avoiding feature loss. Multi-scale convolution captures the fine texture of cracks (shallow layer), the overall outline of cracks (middle layer), and the association information between cracks and surrounding track areas (deep layer) through convolutional kernels of different sizes. The extraction of contextual semantic information can effectively distinguish cracks from similar interfering features such as rust and scratches, improving the accuracy and generalization ability of crack detection. The preset probability threshold is determined as follows: track images containing cracks and non-cracks are collected as a validation set, and the probability distribution map of each image is output through the crack detection network; with a step size of 0.01, the probability thresholds in the range of 0.1 to 0.9 are traversed, each probability distribution map is binarized, the crack detection accuracy and recall rate under each threshold are calculated, and then the F1 score is calculated; the probability value that maximizes the average F1 score of the validation set is selected as the preset probability threshold, and it is pre-stored in the controller for crack candidate region extraction during real-time detection.
[0031] Reference Figure 5 As shown, the generation of detection records containing crack parameters and locations specifically includes: Obtain the crack candidate regions and their pixel-level position coordinates after connected component analysis, where pixel value 1 corresponds to the crack candidate region and pixel value 0 corresponds to the background region. The Zhang-Suen thinning algorithm is used to iteratively thin the binary image of the crack candidate region. It traverses all crack pixels with a pixel value of 1 in the image and determines whether each pixel meets the preset thinning conditions. Pixels that meet the thinning conditions are marked and deleted, while pixels that do not meet the conditions are kept. The traversal and deletion steps are repeated until no pixels are marked or deleted in the current traversal, at which point the iteration terminates. After the iteration terminates, the remaining pixels are connected sequentially according to their pixel-level position coordinates to form a single-pixel-width line, which is the crack center line. At the same time, all pixel-level position coordinates of the crack center line are recorded. Traverse all adjacent pixels along the crack centerline, calculate the Euclidean distance between each pair of adjacent pixels, sum them up to obtain the centerline pixel length, and multiply by the image resolution to obtain the actual length of the crack. For each pixel on the crack centerline, calculate the tangent direction of the centerline at that point, and the normal direction is the direction perpendicular to the tangent direction. Scan along the normal direction to both sides, count the number of consecutive crack pixels, multiply by the image resolution to obtain the crack width at that point, and take the maximum width among all centerline points as the maximum crack width. Obtain the depth value at the corresponding position in the depth matrix, and calculate the arithmetic mean of the depth values of all pixels in the crack candidate region as the average crack depth. The length, width, and depth parameters of the crack are bound to the mileage identifier of the current fused data frame, and a detection record is generated and uploaded to the monitoring platform according to the preset crack level classification rules.
[0032] This can be explained by shrinking the crack region to a centerline with a width of one pixel, which facilitates length statistics and determination of the normal direction. The width is calculated by scanning the boundary of the crack region along the normal direction at each point on the centerline. The depth information is directly read from the depth matrix at the corresponding position. Finally, the geometric parameters and position information of the crack are stored in a structured manner, providing a quantitative basis for track maintenance decisions. The preset refinement conditions are as follows: the pixel to be deleted is a cracked region pixel, the number of pixels belonging to the cracked region in its eight neighboring regions is within a closed range of 2 to 6, and the connectivity of the eight neighboring regions is 1, ensuring that the connectivity of the cracked region is not destroyed after deleting the pixel; at the same time, the boundary pixel quantization judgment condition is met, that is, the pixels above, to the right, and below the current pixel are not simultaneously cracked pixels, and the pixels to the right, below, and to the left are not simultaneously cracked pixels, ensuring that only edge pixels are deleted; The preset crack level classification rules are based on a comprehensive determination of crack length, width, and depth parameters, and are specifically divided into three levels: Level 1 (minor crack): The crack length is less than a preset first length threshold, the width is less than a preset first width threshold, and the depth is less than a preset first depth threshold; Level 2 (General Crack): The crack length is between the first length threshold and the second length threshold, or the width is between the first width threshold and the second width threshold, or the depth is between the first depth threshold and the second depth threshold; Level 3 (Severe Crack): Crack length is greater than or equal to the second length threshold, or width is greater than or equal to the second width threshold, or depth is greater than or equal to the second depth threshold.
[0033] The crack classification comprehensively considers the impact of crack size on track structure safety. Level 1 cracks have a relatively small impact on track operation safety, Level 2 cracks need to be included in the scope of regular monitoring, and Level 3 cracks are high-risk defects that affect train operation safety, requiring immediate warning and maintenance. The first length threshold, second length threshold, first width threshold, second width threshold, first depth threshold, and second depth threshold are preset and stored in the controller according to the monorail track design specifications, operation safety standards, and actual detection accuracy.
[0034] In another preferred embodiment, the dynamic adjustment of the fusion weights of the two feature maps specifically involves: The crack detection network outputs the crack probability distribution map, and outputs the crack recognition confidence scores corresponding to the two-dimensional texture feature map and the three-dimensional morphological feature map respectively, and calculates the ratio of the confidence scores of the two-dimensional texture feature map and the three-dimensional morphological feature map. When the confidence ratio is less than the first preset threshold, the contribution of the three-dimensional depth feature to crack identification is greater than that of the two-dimensional texture feature, and it is determined to be in the nascent stage. The fusion weight of the three-dimensional morphological feature map is increased to the first multiple of the initial weight, and the fusion weight of the two-dimensional texture feature map is decreased to the second multiple of the initial weight. When the confidence ratio is between the first preset threshold and the second preset threshold, the contributions of the two types of features are equal, and it is determined to be the expansion period, keeping the initial fusion weight unchanged. When the confidence ratio is greater than the second preset threshold, the contribution of the two-dimensional texture feature to crack recognition is greater than that of the three-dimensional depth feature, and it is determined to be the penetration period. The fusion weight of the two-dimensional texture feature map is increased to the first multiple of the initial weight, and the fusion weight of the three-dimensional morphological feature map is decreased to the second multiple of the initial weight. The first and second preset thresholds are determined based on the statistical distribution of the confidence ratio on the validation set and the manually labeled development period. Specifically, validation set samples containing cracks at different development stages are collected, the confidence ratio of each sample is calculated, the upper quartile of the confidence ratio of the nascent stage samples is used as the first preset threshold, and the lower quartile of the confidence ratio of the through stage samples is used as the second preset threshold. The first and second multiple ranges refer to the values obtained on the validation set, taking nascent samples as the object, the adjusted fusion weights as the optimization variables, and the crack detection accuracy as the optimization objective. The optimal weight adjustment factor that maximizes the detection accuracy of nascent samples is determined through grid search. The optimal multiple obtained from the search is then expanded into a first multiple range (corresponding to increasing the weight) and a second multiple range (corresponding to decreasing the weight). The first multiple range is 1.2 to 2.0, and the second multiple range is 0.5 to 0.8. The upper limit of the first multiple range and the lower limit of the second multiple range have a product of 1.
[0035] In another preferred embodiment, the connected component analysis of the initial crack segmentation mask, the removal of isolated noise points with an area smaller than a preset area threshold, and the retention of crack candidate regions and their pixel-level position coordinates include: Obtain the initial segmentation mask for the crack. This mask is a binary mask with the same resolution as the multidimensional feature map of the crack. Pixel value 1 corresponds to the suspected crack area and pixel value 0 corresponds to the background area. An eight-neighborhood connected component analysis algorithm is used to traverse all regions with a pixel value of 1 in the initial crack segmentation mask, and mark all interconnected independent regions, with each independent region corresponding to a connected component. The eight-neighbor connected component analysis algorithm scans pixel by pixel starting from the top left corner of the mask. When it encounters an unmarked pixel with a value of 1, it starts a new connected component marking and expands the marking of all connected pixels through eight-neighbor recursion or a stack structure. The eight-neighbor connected component analysis algorithm takes the target pixel as the center and considers its adjacent pixels in the vertical, horizontal, and four diagonal directions as connected pixels. By traversing and marking each pixel, it can achieve accurate identification of all connected components. Calculate the total number of pixels contained in each connected component, and use this number of pixels as the actual area of the connected component; A preset area threshold is used to compare the actual area of each connected component with the preset area threshold. Connected components with an actual area smaller than the preset area threshold (i.e., isolated noise points) are eliminated. Connected components with an actual area greater than or equal to the preset area threshold are retained and identified as crack candidate regions. At the same time, the x and y coordinates of all pixels in each crack candidate region are extracted and their pixel-level position coordinates are recorded. The result containing the crack candidate region and its corresponding pixel-level position coordinates is output. The preset area threshold is determined as follows: on the validation set, the area distribution of all connected components after binarization is statistically analyzed, and a double logarithmic curve of area versus number of connected components is plotted. The area value corresponding to the inflection point of the curve is the boundary point between noise and real cracks, and this value is used as the preset area threshold.
[0036] In another preferred embodiment, the training loss function and morphological constraint evaluation index for constructing the crack detection network further include: The training loss function is a weighted sum of cross-entropy loss and Dice loss, used for backpropagation updates of network parameters. Cross-entropy loss is defined as the classification error between the network's predicted probability and the true label, and is used to penalize the class prediction error for each pixel; The Dice loss is defined as 1 minus the ratio of the overlap between the predicted crack region and the actual crack region, used to penalize the overall shape difference between the predicted region and the actual region; wherein, the predicted crack region is directly obtained from the probability map by binarization with a preset probability threshold during the network training stage, without going through connected component analysis. The morphological constraint evaluation index is composed of a weighted average of connectivity loss and curvature smoothing loss, and is used to evaluate the morphological rationality of the prediction results during the verification phase, assisting in the selection of optimal model parameters. Connectivity loss is defined as the ratio of the number of connected domains in the prediction mask to the total area of the crack region, and is used to evaluate the degree of discontinuity in the crack region. Curvature smoothing loss is defined as the average value of the curvature of each pixel on the crack centerline. The curvature is approximated by second-order difference and is used to evaluate the degree of severe bending of the crack centerline. During training, the morphological constraint evaluation index is used as the early stopping criterion and the basis for model selection: after each training cycle, the morphological constraint evaluation index value on the validation set is calculated; when the index no longer improves for several consecutive training cycles, training is stopped; after training is completed, the model parameters with the optimal morphological constraint evaluation index are selected as the final network parameters, thereby indirectly guiding the network to learn prediction results that conform to crack morphology characteristics. The weighting coefficients of the training loss function and the morphological constraint evaluation index are all hyperparameters, which are determined through grid search optimization on the validation set.
[0037] Furthermore, based on the same inventive concept as the aforementioned real-time track crack detection method based on multidimensional machine vision, this solution proposes a real-time track crack detection system based on multidimensional machine vision, comprising: The data acquisition module is used to acquire robot walking displacement data through an encoder and generate a synchronous trigger signal according to a preset displacement interval, which simultaneously triggers data acquisition from the industrial camera and the line laser contour sensor. The data preprocessing module is used to spatiotemporally align the acquired images of the rail top, rail side, and weld area with the track cross-section depth data to form a fused data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix; perform illumination compensation and noise suppression processing on the two-dimensional unfolded image to obtain a standardized image, and input the standardized image into a semantic segmentation network to extract the effective area of the track surface; within the effective area of the track surface, extract the gray-level gradient distribution features of the two-dimensional unfolded image and the local height abrupt change features of the three-dimensional depth data respectively, and perform pixel-level fusion of the two types of features to construct a multi-dimensional feature map of the crack; The crack recognition module is used to input the multi-dimensional feature map of the crack into the crack detection network, output the probability distribution map of the existence of the crack, and extract the pixel-level position information of the crack candidate region based on the probability threshold; based on the pixel-level position information of the crack candidate region, calculate the actual length, maximum width and average depth of the crack, and associate them with the current mileage information to generate a detection record containing crack parameters and position. The data preprocessing module includes: The fusion data unit is used to align the acquired images of the rail top, rail side and weld area with the rail section depth data in time and space to form a fusion data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix. The effective region unit is used to perform illumination compensation and noise suppression processing on the two-dimensional unfolded image to obtain a standardized image, and input the standardized image into the semantic segmentation network to extract the effective region of the track surface. The feature map unit is used to extract the gray-level gradient distribution features of the two-dimensional unfolded map and the local height change features of the three-dimensional depth data in the effective area of the track surface, respectively, and to fuse the two types of features at the pixel level to construct a multi-dimensional feature map of cracks. The crack detection module includes: A crack identification unit is used to input a multi-dimensional feature map of cracks into a crack detection network, output a probability distribution map of crack existence, and extract pixel-level location information of crack candidate regions based on a probability threshold. A crack recording unit is used to calculate the actual length, maximum width, and average depth of a crack based on the pixel-level location information of a crack candidate region, and associate it with the current mileage information to generate a detection record containing crack parameters and location.
[0038] In summary, the advantages of this invention are: through multi-sensor synchronous acquisition and two-dimensional and three-dimensional feature fusion detection, it can achieve accurate identification and quantitative classification of track cracks, with stable detection, strong anti-interference, and real-time automated inspection.
[0039] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A real-time detection method for track cracks based on multi-dimensional machine vision, characterized in that, include: The robot's walking displacement data is collected by the encoder, and a synchronous trigger signal is generated according to the preset displacement interval, which simultaneously triggers the data acquisition of the industrial camera and the line laser contour sensor. The acquired images of the rail top, rail side and weld area are spatiotemporally aligned with the rail section depth data to form a fused data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix. Illumination compensation and noise suppression are performed on the two-dimensional unfolded image to obtain a standardized image, which is then input into a semantic segmentation network to extract the effective region of the track surface. Within the effective area of the track surface, the gray-level gradient distribution features of the two-dimensional unfolded image and the local height change features of the three-dimensional depth data are extracted respectively. The two types of features are fused at the pixel level to construct a multi-dimensional feature map of the crack. The multidimensional feature map of the crack is input into the crack detection network, which outputs the probability distribution map of the existence of cracks, and extracts the pixel-level location information of the crack candidate region based on the probability threshold. Based on the pixel-level location information of the crack candidate region, the actual length, maximum width and average depth of the crack are calculated and associated with the current mileage information to generate a detection record containing crack parameters and location.
2. The method for real-time detection of track cracks based on multi-dimensional machine vision according to claim 1, characterized in that, The process of acquiring robot displacement data via encoder and generating synchronous trigger signals based on preset displacement intervals to simultaneously trigger data acquisition from industrial cameras and line laser contour sensors specifically includes: Control the monorail inspection robot to move at a constant speed along the monorail track to be tested; The quadrature pulse signal output by the encoder is acquired and converted into a real-time walking displacement value by a counter; When the real-time walking displacement value reaches the preset displacement interval, a trigger pulse is generated and simultaneously output to the industrial cameras and line laser contour sensors arranged in top view, side view, and oblique view. After receiving a trigger pulse, the industrial camera and the line laser profile sensor simultaneously perform a single image acquisition and cross-sectional profile scanning. The real-time walking displacement value corresponding to the trigger time is used as the mileage identifier of the data frame collected this time, and is bound and stored with the collected multi-view images and cross-sectional contour data.
3. The method for real-time detection of track cracks based on multi-dimensional machine vision according to claim 2, characterized in that, The step of spatiotemporally aligning the acquired images of the rail top, rail side, and weld area with the rail cross-section depth data to form a fused data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix specifically includes: The images of the rail top, rail side and weld area at the same trigger moment are registered according to the preset spatial splicing template to generate a two-dimensional unfolded view of the entire rail section. Extract the cross-sectional contour point cloud output by the line laser contour sensor, and generate a depth matrix corresponding to the pixel coordinates of the two-dimensional unfolded image through a linear interpolation algorithm; Using mileage identifiers as the primary key, the two-dimensional unfolded image and depth matrix are encapsulated into a fused data frame to achieve point-to-point correspondence between image pixels and depth values.
4. The method for real-time detection of track cracks based on multi-dimensional machine vision according to claim 3, characterized in that, The process of performing illumination compensation and noise suppression on the two-dimensional unfolded image to obtain a standardized image, and then inputting the standardized image into a semantic segmentation network to extract the effective region of the track surface, specifically includes: A two-dimensional unfolded image is obtained from the fused data frame. Adaptive histogram equalization is used to adjust the local contrast of the image. Bilateral filtering is used to retain edge information while suppressing high-frequency noise to obtain a standardized image. The processed, normalized image is input into a lightweight semantic segmentation network with an encoder-decoder structure, and the output is a binary mask of the track surface region; A binary mask is used to filter pixels in the standardized image, retaining pixels in the orbital region to obtain an image of the effective region on the orbital surface.
5. The method for real-time detection of track cracks based on multi-dimensional machine vision according to claim 4, characterized in that, The step of extracting the gray-level gradient distribution features of the two-dimensional unfolded image and the local height abrupt change features of the three-dimensional depth data within the effective area of the track surface, and fusing the two types of features at the pixel level to construct a multi-dimensional crack feature map specifically includes: Within the effective area of the track surface, a multi-directional Gabor filter bank is used to extract texture responses at different angles to generate a two-dimensional texture feature map; The depth matrix is subjected to first-order difference operation to extract the height change rate along the track direction and perpendicular to the track direction, and the local depth anomaly is calculated. After integration, a three-dimensional morphological feature map is generated. The two-dimensional texture feature map and the three-dimensional morphological feature map are weighted and superimposed according to pixel position to construct a multi-dimensional feature map of cracks.
6. The method for real-time detection of track cracks based on multi-dimensional machine vision according to claim 5, characterized in that, The process of inputting the multidimensional feature map of the crack into the crack detection network, outputting a probability distribution map of the existence of cracks, and extracting pixel-level location information of crack candidate regions based on probability thresholds specifically includes: The crack detection network adopts a fully convolutional structure. It takes the multi-dimensional feature map of cracks as input, extracts contextual semantic information through multi-scale convolution, and after passing through the Sigmoid activation function, outputs a probability map with the same resolution as the input image, as well as the crack recognition confidence of the two-dimensional texture feature map and the three-dimensional morphological feature map. The multi-scale convolution includes a combination of asymmetric convolution kernels, and the value at each pixel position in the probability map represents the probability value that the pixel belongs to the crack category. The crack development period is divided into the initiation stage, the propagation stage, and the penetration stage. Based on the crack recognition confidence of the two feature maps, the fusion weight of the two feature maps is dynamically adjusted. The training loss function and morphological constraint evaluation index of the crack detection network are constructed. The training loss function is a weighted sum of cross-entropy loss and Dice loss. The morphological constraint evaluation index is a weighted sum of connectivity loss and curvature smoothing loss. The probability map is binarized according to a preset probability threshold. Pixels with probability values greater than or equal to the threshold are marked as crack candidate pixels, and pixels with probability values less than the threshold are marked as background pixels, thus obtaining the initial crack segmentation mask. Connectivity analysis is performed on the initial crack segmentation mask to remove isolated noise points with an area smaller than a preset area threshold, while retaining the crack candidate region and its pixel-level position coordinates.
7. The method for real-time detection of track cracks based on multi-dimensional machine vision according to claim 6, characterized in that, The process of calculating the actual length, maximum width, and average depth of the crack based on the pixel-level location information of the crack candidate region, and then associating this information with the current mileage information to generate a detection record containing crack parameters and location specifically includes: Obtain the crack candidate regions and their pixel-level position coordinates after connected component analysis, where pixel value 1 corresponds to the crack candidate region and pixel value 0 corresponds to the background region. The Zhang-Suen thinning algorithm is used to iteratively thin the binary image of the crack candidate region. It traverses all crack pixels with a pixel value of 1 in the image and determines whether each pixel meets the preset thinning conditions. Pixels that meet the thinning conditions are marked and deleted, while pixels that do not meet the conditions are kept. The traversal and deletion steps are repeated until no pixels are marked or deleted in the current traversal, at which point the iteration terminates. After the iteration terminates, the remaining pixels are connected sequentially according to their pixel-level position coordinates to form a single-pixel-width line, which is the crack center line. At the same time, all pixel-level position coordinates of the crack center line are recorded. Traverse all adjacent pixels along the crack centerline, calculate the Euclidean distance between each pair of adjacent pixels, sum them up to obtain the centerline pixel length, and multiply by the image resolution to obtain the actual length of the crack. For each pixel on the crack centerline, calculate the tangent direction of the centerline at that point, and the normal direction is the direction perpendicular to the tangent direction. Scan along the normal direction to both sides, count the number of consecutive crack pixels, multiply by the image resolution to obtain the crack width at that point, and take the maximum width among all centerline points as the maximum crack width. Obtain the depth value at the corresponding position in the depth matrix, and calculate the arithmetic mean of the depth values of all pixels in the crack candidate region as the average crack depth. The length, width, and depth parameters of the crack are bound to the mileage identifier of the current fused data frame, and a detection record is generated and uploaded to the monitoring platform according to the preset crack level classification rules.
8. A real-time track crack detection system based on multi-dimensional machine vision, characterized in that, The method for real-time detection of track cracks based on multidimensional machine vision as described in any one of claims 1-7 includes: The data acquisition module is used to acquire robot walking displacement data through an encoder and generate a synchronous trigger signal according to a preset displacement interval, which simultaneously triggers data acquisition from the industrial camera and the line laser contour sensor. The data preprocessing module is used to spatiotemporally align the acquired images of the rail top, rail side, and weld area with the track cross-section depth data to form a fused data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix; perform illumination compensation and noise suppression processing on the two-dimensional unfolded image to obtain a standardized image, and input the standardized image into a semantic segmentation network to extract the effective area of the track surface; within the effective area of the track surface, extract the gray-level gradient distribution features of the two-dimensional unfolded image and the local height abrupt change features of the three-dimensional depth data respectively, and perform pixel-level fusion of the two types of features to construct a multi-dimensional feature map of the crack; The crack recognition module is used to input the multi-dimensional feature map of the crack into the crack detection network, output the probability distribution map of the existence of the crack, and extract the pixel-level position information of the crack candidate region based on the probability threshold; based on the pixel-level position information of the crack candidate region, calculate the actual length, maximum width and average depth of the crack, and associate them with the current mileage information to generate a detection record containing crack parameters and position.
9. A real-time track crack detection system based on multi-dimensional machine vision according to claim 8, characterized in that, The data preprocessing module includes: The fusion data unit is used to align the acquired images of the rail top, rail side and weld area with the rail section depth data in time and space to form a fusion data frame containing a two-dimensional unfolded image and a three-dimensional depth matrix. The effective region unit is used to perform illumination compensation and noise suppression processing on the two-dimensional unfolded image to obtain a standardized image, and input the standardized image into the semantic segmentation network to extract the effective region of the track surface. The feature map unit is used to extract the gray-level gradient distribution features of the two-dimensional unfolded map and the local height change features of the three-dimensional depth data within the effective area of the track surface, respectively, and to fuse the two types of features at the pixel level to construct a multi-dimensional feature map of the crack.
10. A real-time track crack detection system based on multi-dimensional machine vision according to claim 9, characterized in that, The crack detection module includes: A crack identification unit is used to input a multi-dimensional feature map of cracks into a crack detection network, output a probability distribution map of crack existence, and extract pixel-level location information of crack candidate regions based on a probability threshold. A crack recording unit is used to calculate the actual length, maximum width, and average depth of a crack based on the pixel-level location information of a crack candidate region, and associate it with the current mileage information to generate a detection record containing crack parameters and location.