Station structure damage identification method and system based on image recognition

By using image recognition technology to calculate the apparent magnification factor and correct the image sequence, the problem of false magnification of minute cracks in images during train operation was solved, enabling more accurate crack width measurement and safety warning.

CN122289208APending Publication Date: 2026-06-26GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-31
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively address the problem of false magnification of the apparent width of minute cracks in images caused by the vibration transmission effect of the station's load-bearing system during train operation and the interweaving of the line-by-line exposure sequence of the camera equipment, which affects the objective accuracy of damage assessment and the reliability of early warning.

Method used

By acquiring image sequences, the train operation period is extracted using the grayscale changes in the reference area. The apparent magnification factor is calculated and the image is corrected. The artificially inflated width is deducted, and the actual crack width and spalling area are converted. The warning level is calculated in combination with the importance of the load-bearing component.

Benefits of technology

The system accurately deducts the artificially increased width compensation item caused by the station scene, improving the objective accuracy of damage identification of the station load-bearing system and the reliability of safety warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289208A_ABST
    Figure CN122289208A_ABST
Patent Text Reader

Abstract

This invention relates to the field of image processing technology and discloses a method and system for identifying station structural damage based on image recognition. The method includes: extracting the train's operating period using grayscale changes in a reference area, and obtaining imaging displacement and principal vibration parameters by tracking non-damaged parts of the reference area; calculating the apparent magnification factor by combining spatial scale, vibration direction, amplitude, frequency, and camera exposure delay; using this factor as a driving parameter to perform row-level correction and fusion of the image sequence; and after extracting candidate cracks and spalling areas, accurately subtracting the artificially inflated width from the pixel width to calculate the actual crack width and actual spalling area; finally, calculating a score based on a preset importance and outputting the warning level. This invention can effectively eliminate the false widening error caused by train vibration and prevent fluctuations and artificially high safety warning levels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to a method and system for identifying station structural damage based on image recognition. Background Technology

[0002] With the popularization of machine vision technology, the continuous monitoring of surface damage such as cracks and spalling in station load-bearing systems (such as central columns, side walls, and nodes) using fixed high-definition cameras has gradually become an important non-contact inspection method. Currently, most high-definition cameras in monitoring scenarios acquire image sequences using a line-by-line exposure method. In actual service environments, the station load-bearing system inevitably experiences vibration transmission effects due to train loads during train entry and exit and during near- and far-rail overlap. Existing visual inspection solutions typically treat station vibration transmission, line-by-line exposure degradation, and crack normal width measurement as isolated components for separate processing or avoidance. However, during train operation, the vibration displacement of the load-bearing system intertwines with the line-by-line exposure sequence of the camera, easily leading to a false magnification of the apparent width of minute cracks during the imaging stage. Existing technologies have not achieved in-depth integration and explicit modeling of the above multiple environmental factors, resulting in a systematic overestimation in the calibration conversion of pixel width to actual spatial width when processing images from train operation periods. This artificially inflated error, which has not yet been effectively identified and deducted, directly disrupts the judgment link between calibration parameters and the real state. As a result, the final output safety warning level frequently shows artificially high values ​​and fluctuates when the train passes by, which seriously restricts the objective accuracy of damage assessment and the reliability of warning. Summary of the Invention

[0003] This invention provides a method and system for identifying station structural damage based on image recognition, which solves the technical problems mentioned in the background art.

[0004] The first aspect is a station structural damage identification system based on image recognition, including: The acquisition and extraction module acquires image sequences of the carrier through camera equipment, extracts the train operation period using grayscale changes in a preset reference area, and converts the pixel scale into a spatial scale. The vibration estimation module obtains the imaging displacement by tracking the non-damaged parts of the reference area during the train's operation period, and calculates the vibration direction, amplitude, frequency, and initial phase. The restoration module obtains candidate cracks from the image sequence, extracts the normal direction and span value of the candidate cracks, and calculates the apparent magnification factor by combining the spatial scale, vibration direction, amplitude, frequency, initial phase and exposure delay of the camera device. The image sequence is then corrected using the apparent magnification factor to obtain a corrected image. The detection module extracts the candidate cracks and peeling areas from the corrected image and measures the pixel width of the candidate cracks in the normal direction. The compensation module calculates the actual crack width and the actual peeling area by subtracting the artificially increased width from the pixel width based on the apparent magnification factor. The early warning module summarizes the actual crack width and the actual spalling area, calculates a score based on the preset importance of the load-bearing component, and outputs the early warning level.

[0005] Secondly, the image recognition-based station structure damage identification method, applied to any of the image recognition-based station structure damage identification systems described above, includes: The image sequence of the carrier is acquired by the camera equipment, the train operation period is extracted by the gray scale change of the preset reference area, and the pixel scale is converted into the spatial scale. During the train's operation, the imaging displacement is obtained by tracking the non-damaged parts of the reference area, and the vibration direction, amplitude, frequency, and initial phase are determined. Candidate cracks are obtained from the image sequence, and the normal direction and span value of the candidate cracks are extracted. The apparent magnification factor is calculated by combining the spatial scale, vibration direction, amplitude, frequency, initial phase and exposure delay of the camera device. The image sequence is then corrected using the apparent magnification factor to obtain a corrected image. The candidate cracks and peeling areas are extracted from the corrected image, and the pixel width of the candidate cracks is measured in the normal direction. Based on the apparent magnification factor, the artificially increased width is subtracted from the pixel width to calculate the actual crack width and the actual peeling area. The actual crack width and the actual spalling area are combined, and a score is calculated based on the preset importance of the load-bearing component, and an early warning level is output.

[0006] The beneficial effects of this invention are as follows: By establishing a station phase-locked apparent crack width magnification coefficient as a core parameter, the system deeply integrates the vibration transmission effect of the load-bearing system caused by train operation, the line-by-line exposure sequence of the camera equipment, and the crack direction, thus completely solving the imaging pseudo-widening problem caused by the interplay of these multiple factors. Before converting the pixel width into a spatial scale mapping, the system accurately deducts the illusory width compensation term caused by the station-specific scene, thereby avoiding false height warnings triggered by imaging distortion of unexpanded cracks. At the same time, this invention combines the multi-dimensional actual crack scale, actual spalling area, and preset importance of different locations in the load-bearing system to achieve more accurate graded risk output, significantly improving the objective accuracy of damage identification of the station load-bearing system and the reliability of safety warnings. Attached Figure Description

[0007] Figure 1 This is a flowchart illustrating the implementation of the image recognition-based station structure damage identification system of the present invention. Figure 2 This is a schematic diagram of a specific implementation scenario of the present invention. Detailed Implementation

[0008] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.

[0009] Example 1: As Figure 1 As shown, the image recognition-based station structural damage identification system includes: The acquisition and extraction module acquires image sequences of the carrier through camera equipment, extracts the train operation period using grayscale changes in a preset reference area, and converts the pixel scale into a spatial scale. The vibration estimation module obtains the imaging displacement by tracking the non-damaged parts of the reference area during the train's operation period, and calculates the vibration direction, amplitude, frequency, and initial phase. The restoration module obtains candidate cracks from the image sequence, extracts the normal direction and span value of the candidate cracks, and calculates the apparent magnification factor by combining the spatial scale, vibration direction, amplitude, frequency, initial phase and exposure delay of the camera device. The image sequence is then corrected using the apparent magnification factor to obtain a corrected image. The detection module extracts the candidate cracks and peeling areas from the corrected image and measures the pixel width of the candidate cracks in the normal direction. The compensation module calculates the actual crack width and the actual peeling area by subtracting the artificially increased width from the pixel width based on the apparent magnification factor. The early warning module summarizes the actual crack width and the actual spalling area, calculates a score based on the preset importance of the load-bearing component, and outputs the early warning level.

[0010] Preferably, the process involves acquiring image sequences of the carrier component using a camera device, extracting the train's operating period using grayscale changes in a reference area, and converting the pixel scale into a spatial scale, including: The current frame image and the previous frame image are obtained from the image sequence. The average value of the absolute value of the gray level difference between all pixels in the reference region and the current frame image is calculated to obtain the inter-frame energy mean. The calculation formula is as follows:

[0011] in, This represents the average energy value between frames; This represents the set of pixels within the reference region; This represents the total number of pixels within the reference area; Represents pixel coordinates; Indicates the coordinates of the current frame image The grayscale value at that location; This indicates that the previous frame image is located at coordinates The grayscale value at that location; Obtain the mean and standard deviation of the baseline energy of the reference region under historical stable conditions, and calculate the operational period judgment threshold; the calculation formula is as follows:

[0012] in, This indicates the runtime determination threshold; This represents the baseline energy mean; This represents the baseline energy standard deviation; The set of consecutive image frames whose average inter-frame energy is greater than the operating period determination threshold is determined as the train operating period; Obtain the known reference length within the reference region and the corresponding image pixel span, and calculate the spatial scale; the calculation formula is as follows:

[0013] in, This represents the converted spatial scale; Indicates the known reference length; This indicates the pixel span of the image.

[0014] The set of pixels within the reference region is a pre-defined set of pixel regions in the image of the load-bearing component, used for extraction and vibration parameter inversion during train operation. Preferably, each component region is divided into 5 pixel blocks of 64 by 64 pixels. This size of pixel block contains sufficient texture details for grayscale variation analysis, while avoiding the reduction in computational efficiency caused by excessively large regions. The number of 5 blocks ensures the robustness of the analysis and prevents texture anomalies in a single region from affecting the overall analysis results.

[0015] The grayscale value at coordinates of the current frame image is the normalized grayscale value of the nth frame image in the image sequence captured by the camera device at a specified pixel coordinate position. It can be obtained by performing grayscale and normalization processing on the image after it has been captured by the camera device using digital image processing algorithms. The commonly used grayscale algorithm is the weighted average method, and the normalization process can map the grayscale value to a value range of 0 to 255.

[0016] The grayscale value at coordinates of the previous frame is the normalized grayscale value of the n-1th frame in the image sequence captured by the camera at a specified pixel coordinate position. It can be obtained by extracting the previous frame from the continuous image sequence captured by the camera, and then performing grayscale and normalization processing using digital image processing algorithms. The grayscale and normalization processing methods are completely consistent with those of the current frame.

[0017] The baseline energy mean is the statistical average of the inter-frame energy mean within the reference area under historical stable conditions, where the stable conditions refer to the normal acquisition state of the camera equipment when no trains pass through the station.

[0018] The baseline energy standard deviation is the statistical standard deviation of the inter-frame energy mean in the reference area under historical steady conditions. It is used to characterize the dispersion of the inter-frame energy mean under steady conditions. The larger the value, the more obvious the fluctuation of the inter-frame grayscale change under steady conditions.

[0019] The train operation period determination threshold is the critical value of the inter-frame average energy used to distinguish whether a train has passed through a station, and it is the core threshold for determining the train operation period.

[0020] The inter-frame energy mean is the arithmetic mean of the absolute values ​​of the gray level difference between the current frame and the previous frame for all pixels in the reference region. It is used to characterize the degree of gray level change between image frames. The larger the value, the more drastic the gray level change between frames.

[0021] The known reference length within the reference area is the length of a line segment with actual physical dimensions that has been pre-marked within the reference area. Preferably, line segments with fixed physical dimensions, such as structural joints or embedded markers within the reference area, are selected. The preferred length is 50 mm to 200 mm. Line segments within this length range can form a suitable pixel span in the image, facilitating accurate measurement of the corresponding pixel distance and effectively improving the accuracy of spatial scale conversion.

[0022] The pixel span corresponding to a known reference length is the number of pixels in the image corresponding to the known reference length within the reference region, i.e., the pixel distance of the line segment along its length in the image. This can be obtained through pixel counting or line fitting distance measurement methods in digital image processing algorithms. First, edge detection and line fitting are performed on the reference length line segment in the image. Then, the number of pixels on the fitted line or the pixel distance between the two endpoints of the line segment is calculated.

[0023] Spatial scale is a scaling factor that converts pixel distances in an image into actual physical distances, measured in millimeters per pixel. It is calculated by dividing a known reference length by the corresponding image pixel span.

[0024] The total number of pixels within the reference region is the total number of pixels contained in the predefined reference region, calculated from the size of the reference region and the pixel resolution of the image.

[0025] Image pixel coordinates are two-dimensional coordinates used to locate the position of each pixel in an image. The horizontal and vertical coordinates represent the column and row positions of the pixel in the image, respectively.

[0026] In detail, the average absolute value of the grayscale difference in the reference area is used as the inter-frame energy mean, which is then used as the basis for determining the train operation period. Essentially, this converts the station structural vibrations caused by train passage into intrinsic grayscale energy features within the image. It does not rely on external equipment such as train timetables or vibration sensors, but relies entirely on image data collected by the camera equipment for determination, unlike conventional security videos which use common triggering logic such as inter-frame pixel change rate. The threshold for determining the train operation period is set by adding twice the baseline energy standard deviation to the historical stable baseline energy mean. This threshold setting method is designed based on the actual statistical characteristics of station operation data. The value of the standard deviation can effectively avoid misjudgment caused by random gray-level fluctuations in a stable state, and is suitable for the characteristic of continuous and significant gray-level changes when a train passes through. The set of continuous image frames with an inter-frame energy mean greater than the judgment threshold is defined as the train operation period. The judgment condition of continuous frames is emphasized because the train passing through the station is a continuous process. The gray-level abrupt change in a single frame may be caused by environmental interference. For example, three consecutive frames are selected as the judgment standard. When the inter-frame energy mean of three consecutive frames is greater than the judgment threshold, it can be determined that the train has entered the operation period. This judgment method can greatly improve the accuracy and robustness of the train operation period extraction and is suitable for the actual scene characteristics of station operation.

[0027] In detail, the specific selection rules for the reference area are as follows: Five 64x64 pixel blocks are defined for each station component area. The selection must meet the following conditions: no overlap with historical damage masks; the linear response of cracks must be below the top 10% of the historical distribution in that area; the texture variance must be between 30% and 70% of the historical distribution in that area; and the distance from the area boundary must be at least 20 pixels. For example, for the central cylindrical component area, one reference pixel block can be defined at each of the five positions (top, bottom, left, right, and center) of the cylinder, avoiding already detected crack areas. The definition standard for a historical stable state is to select 1000 to 2000 frames of images continuously captured by the camera when no trains are passing through the station as a stable baseline. During the selection process, the coefficient of variation of the mean energy between frames must be less than 0.1. The coefficient of variation is the ratio of the baseline energy standard deviation to the baseline energy mean. This standard ensures that the selected baseline represents a truly stable state without vibration interference. The specific number of frames in the continuous image during train operation is determined. The system uses a 3-frame rule: the start of a train operation period is determined when the average inter-frame energy of three consecutive frames exceeds the threshold for determining the train operation period; the end of a train operation period is determined when the average inter-frame energy of ten consecutive frames is less than or equal to the average baseline energy plus one time the baseline energy standard deviation. This frame count is determined by combining the time characteristics of trains passing through stations with the frame rate characteristics of image acquisition. For example, when the camera equipment has a frame rate of 30 frames per second, a 3-frame rule can quickly identify the train's entry, while a 10-frame rule can avoid the mistaken end of the operation period caused by temporary train deceleration. The specific calibration method for the known reference length within the reference area is to select the inherent fixed dimensional characteristics of the station structure within the reference area, such as the joints of concrete components, embedded metal marking strips, and equidistant scales reserved in the structure. A high-precision laser rangefinder is used to measure their actual physical length. The measurement accuracy of the laser rangefinder is not less than 0.1 mm, ensuring the accuracy of the known reference length and providing a reliable benchmark for spatial scale conversion.

[0028] Preferably, during the train's operation, the imaging displacement is obtained by tracking the non-damaged portion of the reference area, and the vibration direction, amplitude, frequency, and initial phase are determined, including: Multiple reference tracking points are set at the non-damaged portion of the reference region, and the two-dimensional imaging displacement vector of each reference tracking point in the image frame is obtained; the calculation formula is as follows:

[0029] in, Indicates the first The reference tracking point at the first The two-dimensional imaging displacement vector of each image frame; Indicates the horizontal displacement component; Represents the vertical displacement component; Calculate the time-mean displacement vector of each of the aforementioned reference tracking points and the overall mean displacement vector of all the aforementioned reference tracking points, and then calculate the displacement covariance matrix; the calculation formula is as follows:

[0030] in, Represents the displacement covariance matrix; This indicates the total number of the reference tracking points; Indicates the first The time-averaged displacement vector of each of the reference tracking points; This represents the overall mean displacement vector of all the aforementioned reference tracking points; The eigenvector corresponding to the largest eigenvalue of the displacement covariance matrix is ​​extracted, and the direction of the eigenvector is determined as the vibration direction; the calculation formula is as follows:

[0031] in, This represents the eigenvector corresponding to the largest eigenvalue; Indicates the direction of the vibration; The two-dimensional imaging displacement vector of each reference tracking point is projected onto the vibration direction to obtain the projected displacement; the calculation formula is as follows:

[0032] in, Indicates the first The aforementioned reference tracking points at the corresponding time The projected displacement; By fitting a sine function to the projected displacement of each reference tracking point, the single-point vibration amplitude, single-point vibration frequency, and single-point initial phase of each reference tracking point are obtained; the calculation formula is as follows:

[0033] in, Indicates the first The single-point vibration amplitude of the reference tracking points; Indicates the first The single-point vibration frequency of the reference tracking points; Indicates the first The initial phase of each of the aforementioned reference tracking points; This represents the displacement offset term corresponding to the tracking point; Based on preset tracking point weights, a weighted average is calculated for the single-point vibration amplitude, single-point vibration frequency, and single-point initial phase of each reference tracking point to obtain the amplitude, frequency, and initial phase, respectively; the calculation formula is as follows:

[0034] in, Indicates the amplitude; Indicates the frequency; Indicates the initial phase; Indicates the first The tracking point weights of the reference tracking points.

[0035] The total number of reference tracking points is the total number of feature points set in the non-damaged parts of the reference area for tracking imaging displacement. Preferably, 20 to 30 reference tracking points are set in each reference area. This range of numbers ensures a sufficient sample size for displacement tracking, improving the accuracy of vibration parameter calculation, while preventing a surge in computational load due to an excessive number of points, thus ensuring the real-time processing efficiency of the system.

[0036] The two-dimensional imaging displacement vector is a two-dimensional vector consisting of the horizontal and vertical displacements of the p-th reference tracking point relative to its initial position in the n-th image frame.

[0037] The horizontal displacement component is the displacement value of the reference tracking point in the horizontal direction of the image in the two-dimensional imaging displacement vector, and the unit is pixels.

[0038] The vertical displacement component is the displacement value of the reference tracking point in the vertical direction of the image in the two-dimensional imaging displacement vector, and the unit is pixels.

[0039] The time-averaged displacement vector is the statistical average of the two-dimensional imaging displacement vectors of the p-th reference tracking point across all image frames during the train's operation, used to characterize the overall displacement trend of that reference point.

[0040] The overall mean displacement vector is the statistical average of the time mean displacement vectors of all reference tracking points, used to characterize the overall displacement trend of all tracking points within the reference area.

[0041] The displacement covariance matrix is ​​a second-order square matrix calculated from the displacement deviations of all reference tracking points. It is used to characterize the correlation between the horizontal and vertical displacement components of the reference tracking points.

[0042] The eigenvector corresponding to the largest eigenvalue is the eigenvector corresponding to the largest eigenvalue after the displacement covariance matrix is ​​decomposed. The direction of this vector is consistent with the main vibration direction of the station structure.

[0043] The vibration direction is the main vibration direction of the station structure during train operation. It is determined by the direction of the eigenvector corresponding to the largest eigenvalue of the displacement covariance matrix, and the unit is radians.

[0044] The acquisition time is the specific point in time when the nth image frame is captured by the camera device. It can be extracted from the timestamp information automatically added by the camera device when acquiring the image, or it can be calculated from the frame rate and frame number of the image acquisition.

[0045] Projected displacement is a one-dimensional displacement value obtained by projecting the two-dimensional imaging displacement vector of the reference tracking point onto the direction of structural vibration, and the unit is pixels.

[0046] The single-point vibration amplitude is the amplitude of the sine curve of the projected displacement of the p-th reference tracking point as a function of time, expressed in pixels, and represents the magnitude of the vibration amplitude at that reference point.

[0047] The single-point vibration frequency is the frequency of the sinusoidal curve showing the projected displacement of the p-th reference tracking point as a function of time, measured in Hertz, and characterizes the rate of vibration at that reference point. The initial phase of a single point is the initial phase of the sinusoidal curve of the projected displacement of the p-th reference tracking point as a function of time, expressed in radians, and characterizes the initial state of vibration of that reference point.

[0048] The displacement bias term is a constant term in the sinusoidal fitting formula of the projected displacement of the p-th reference tracking point, with units of pixels, and is used to eliminate systematic bias in the displacement fitting.

[0049] The preset weights are weight values ​​set for each reference tracking point and used for weighted average calculation of regional vibration parameters. The values ​​are positively correlated with the displacement tracking stability of the reference point. Preferably, the weights are assigned based on the displacement tracking determination coefficient of the reference point: points with a tracking determination coefficient greater than or equal to 0.95 have a weight of 1.0, points between 0.9 and 0.95 have a weight of 0.8, and points less than 0.9 have a weight of 0.5. The rationale is to ensure that reference points with more stable displacement tracking have a higher weight in the regional parameter calculation, reducing the interference of unstable points on the overall parameters and improving the accuracy of the vibration parameters.

[0050] The overall vibration amplitude of the region is a value obtained by weighted averaging of the single-point vibration amplitudes of all reference tracking points, with the unit being pixels, representing the overall vibration amplitude of the structure within the reference region.

[0051] The overall dominant frequency of the region is a value obtained by weighted averaging of the single-point vibration frequencies of all reference tracking points, and the unit is Hertz. It represents the overall vibration frequency of the structure within the reference region.

[0052] The overall initial phase of vibration in the region is a value obtained by weighted averaging of the initial phases of all reference tracking points, in radians, representing the overall initial state of structural vibration within the reference region.

[0053] In detail, using undamaged areas of the reference region as vibration transmission probes leverages the characteristic that station structure vibrations cause changes in the imaging displacement of reference points. By tracking the displacement of reference points in the undamaged area, the true vibration parameters of the structure are inverted, rather than treating the undamaged area merely as a meaningless image background. For example, selecting points in the flat, undamaged concrete area of ​​the station's central column and tracking the displacement changes of these points can reflect the column's vibration state. This differs from traditional visual inspection methods that only focus on damaged areas and ignore vibration information in the background. The dominant vibration direction is determined by using the eigenvector corresponding to the maximum eigenvalue of the displacement covariance matrix because station structures affected by train vibrations often exhibit unidirectional dominant vibration characteristics. The method reduces two-dimensional displacement to one-dimensional projected displacement, significantly reducing the complexity of subsequent vibration parameter fitting while retaining core vibration information. It extracts single-point vibration parameters by fitting a sinusoidal function to the projected displacement, and then obtains the overall regional parameters through a weighted average with preset weights, rather than using a simple arithmetic average. This is because the displacement tracking stability of different reference points varies, and the arithmetic average would be affected by unstable points. For example, points with poor tracking stability have lower weights, while points with good stability have higher weights. This allows the calculated regional vibration parameters to better reflect the actual vibration state of the structure. This method is specifically designed for the vibration characteristics of station structures and the actual situation of image displacement tracking, and differs from general displacement parameter statistical methods.

[0054] In detail, the selection rules for reference tracking points in the non-damaged areas of the reference region are as follows: they are uniformly distributed within the reference region, with a pixel spacing of 8 to 10 pixels between adjacent points. The selected points must be corner points or texture feature points in the image, avoiding areas with blurred textures and uniform pixel values. The imaging displacement tracking method for the reference tracking points uses corner detection combined with optical flow. First, feature points within the reference region are extracted using a corner detection algorithm. Then, a pyramid optical flow method is used to track the positional changes of these feature points in consecutive frames, calculating the displacement value relative to the initial position. The displacement covariance matrix is ​​calculated to retain 6 decimal places to ensure the accuracy of the principal vibration direction extraction after eigenvalue decomposition. The least squares method is used to fit the projected displacement to a sine function, and... The coefficient of determination for the fit is required to be no less than 0.9. Reference points with a coefficient of determination lower than 0.9 are considered to have unstable tracking, and their preset weights are reduced. When the overlap between the train's near and far rails causes vibration aliasing, the event window during train operation is divided into short sub-windows with a length of 0.4 seconds and an overlap rate of 50%. Vibration parameters are fitted for each sub-window, and the coefficient of determination is calculated. The vibration parameters corresponding to the sub-window with the largest coefficient of determination are selected as the overall vibration parameters of the reference area to ensure the stability of the vibration parameters. The preset weights of the reference tracking points are assigned based on the coefficient of determination for displacement tracking. First, the coefficient of determination for the sinusoidal fitting of the projected displacement of each point is calculated, and then the weights are assigned according to the coefficient of determination interval. Points without a unified assignment are directly calculated using the arithmetic mean.

[0055] Preferably, the normal direction and span value of the candidate cracks are extracted, and the apparent magnification factor is calculated by combining the spatial scale, the vibration direction, the amplitude, the frequency, the initial phase, and the exposure delay of the camera device, including: Edge extraction is performed on the fused image to obtain the candidate cracks. The number of sampling points of the candidate cracks and the normal pixel width of each sampling point are obtained. The average value of all the normal pixel widths is multiplied by the spatial scale to obtain the average width of the candidate cracks. The calculation formula is as follows:

[0056] in, This indicates that the candidate cracks have a uniform width; Indicates the spatial scale; Indicates the number of sampling points; Indicates the first The normal pixel width of each sampling point; The camera exposure time of the camera device is obtained, and the apparent magnification factor is calculated by combining the spatial scale, vibration direction, normal direction, amplitude, frequency, span value, and exposure delay; the calculation formula is as follows:

[0057] in, This represents the apparent magnification factor; Indicates the amplitude; Indicates the direction of the vibration; Indicates the normal direction; Indicates the frequency; This indicates the camera exposure time; This indicates the value spanning multiple rows; This indicates the exposure delay; This represents the zero constant.

[0058] The number of sampling points for candidate cracks is the total number of sampling points set on a single candidate crack to statistically measure the normal pixel width of the candidate cracks. Preferably, one sampling point is set per millimeter of physical crack length. If the physical crack length is less than 10 millimeters, then 10 sampling points are fixed. This method ensures that the sampling point density is suitable for the crack width measurement accuracy. Fixing the number of sampling points for short cracks can avoid the error in average width calculation caused by too few points.

[0059] The normal pixel width is the pixel distance along the crack normal at the m-th sampling point of the k-th candidate crack, and is the core parameter characterizing the crack pixel width at the sampling point. It can be obtained by performing skeletonization processing on the crack mask, drawing a perpendicular line along the crack normal at the sampling point, and counting the number of pixels between the intersection of the perpendicular line and the edge of the crack mask.

[0060] The average width of a candidate crack is the value obtained by multiplying the average normal pixel width of all sampling points of the k-th candidate crack by the spatial scale. The unit is millimeters, and it represents the average physical width of the candidate crack.

[0061] The apparent magnification factor is a dimensionless coefficient calculated by coupling the vibration parameters of the station structure, the exposure parameters of the camera equipment, and the parameters of the crack itself. It represents the degree of magnification of the apparent width of the crack relative to its true width.

[0062] The crack normal direction is the normal angle of the k-th candidate crack in the image, in radians. It is perpendicular to the crack tangential direction and is the core parameter for calculating the vibration displacement along the crack normal component.

[0063] Video exposure time is the duration of exposure of the photosensitive element when a camera captures a single frame of image, measured in seconds. It is an inherent parameter of the camera device. It can be obtained by directly reading the camera's factory settings or configuration file. If the parameter is missing, it can be obtained through the LED panel calibration method, which involves capturing images using a flashing panel with a known frequency and calculating the exposure time.

[0064] The crack row span value is the total number of pixel rows spanned by the k-th candidate crack in the image, representing the vertical extension scale of the crack in the image. It can be obtained by statistically analyzing the difference between the maximum and minimum row numbers corresponding to the crack mask in the image.

[0065] Line exposure delay is the time difference between the start of exposure of two adjacent rows of pixels when a camera uses line-by-line exposure mode. It is measured in seconds per row and is an inherent parameter of the camera. It can be obtained by directly reading the camera's factory settings or configuration file. If the parameter is missing, it can be calculated using the LED panel calibration method: capture an image of a flickering LED panel, extract the tilt angle of the bright stripes, and calculate the line exposure delay.

[0066] The first zero-prevention constant is a very small constant added to the denominator to avoid the denominator being zero when calculating the apparent magnification factor. The unit is millimeters. It is preferably 0.01 millimeters. This value is much smaller than the lower limit of the detection accuracy for fine cracks in the station. Adding it will not affect the calculation result of the apparent magnification factor, and can effectively avoid the problem of the denominator being zero.

[0067] The absolute value of the first sine is the absolute value of the sine function of the product of pi, the dominant frequency of the region as a whole, and the camera exposure time.

[0068] The absolute value of the second sine is the absolute value of the sine function of the product of pi, the overall dominant frequency of the region, the crack span value, and the line exposure delay.

[0069] The magnification numerator is a value obtained by multiplying the constant 4, the spatial scale, the overall vibration amplitude of the region, the absolute value of the cosine of the difference between the vibration direction and the crack normal direction, the absolute value of the first sine, and the absolute value of the second sine. The unit is millimeters, and it is the numerator of the apparent magnification factor.

[0070] The magnification denominator is a value obtained by adding the average width of the candidate crack to the first zero constant, and the unit is millimeters. It is the denominator part for calculating the apparent magnification factor.

[0071] In detail, this method couples the spatial scale related to station structural vibration, the overall vibration amplitude, vibration direction, and overall dominant frequency of the area, the camera exposure time and line exposure delay related to camera exposure, and the crack's own crack normal direction, crack span value, and normal pixel width depth to construct a station-specific apparent magnification factor. This factor can accurately quantify the degree of crack pseudo-widening caused by the coupling of train vibration and line-by-line exposure, which is different from the traditional crack detection method that does not have a scene-specific magnification factor and simply converts pixels and physical width. The average width of the candidate crack is obtained by multiplying the average normal pixel width of all sampling points of the candidate crack by the spatial scale, rather than using the pixel width of a single point to represent the overall crack width. This adapts to the characteristic that the width of fine cracks in the station has local fluctuations and improves the accuracy of the average width calculation. The first zero-prevention constant introduced in the denominator is designed to adapt to the scenario where the average width of fine cracks in the station is close to zero. Its value is much smaller than the detection accuracy limit of fine cracks in the station, which avoids the problem of the denominator being zero and does not have a substantial impact on the calculation result of the apparent magnification factor, which is different from the method of arbitrarily selecting constants in general zero-prevention processing.

[0072] In detail, edge extraction of the initial image frames of the image sequence was performed using the Canney edge detection algorithm, with a low threshold of 50 and a high threshold of 150. This algorithm can effectively suppress noise and accurately extract fine edges of cracks. The selection rule for candidate crack sampling points is to distribute them evenly along the crack skeleton, with a sampling point spacing of 2 to 3 pixels. If the number of pixels in the crack skeleton is less than the number of sampling points, a sampling point is set at each skeleton point. The crack normal direction is calculated by performing a 5-point local linear fitting on the crack skeleton to obtain the tangential angle at each point, adding 90 degrees to the tangential angle to obtain the local normal angle, and then taking the circular average of all local normal angles as the normal direction of the entire crack. The crack span... The statistical rule for row values ​​is based on the pixel row number of the crack mask. The difference between the maximum and minimum row numbers of all pixels in the mask is calculated. If the difference is zero, it is fixed at 1 as the row span value. When the camera does not provide the camera exposure time and row exposure delay, the LED panel calibration method is used. The LED panel is placed in the center of the camera's field of view, the panel flashing frequency is set to 100 Hz, and after taking two frames of images, the bright stripe tilt angle is extracted. The exposure parameters are calculated based on the number of rows spanned and the phase difference. The first zero-prevention constant is fixed at 0.01 mm. It does not need to be adjusted according to the station component type or crack characteristics. This value can be adapted to the detection scenarios of fine cracks in all stations.

[0073] Preferably, correcting the image sequence using the apparent magnification factor to obtain a corrected image includes: Calculate the average of all the aforementioned apparent magnification factors as the mean magnification of the region; the calculation formula is as follows:

[0074] in, This represents the magnified mean value of the region. This represents the total number of apparent magnification factors; Indicates the apparent magnification factor for each of the aforementioned factors; The regional correction intensity is obtained by dividing the mean value of the magnified region by a constant and summing the mean value; the calculation formula is as follows:

[0075] in, Indicates the correction intensity of the area; The product of the exposure delay and the processing row number of the processed image frame is calculated and added to the processing time to obtain the time variable; the constant 2, pi, the frequency, and the time variable are multiplied together, added to the initial phase, and the sine value is taken, then multiplied by the amplitude to obtain the row-level equivalent displacement; the calculation formula is as follows:

[0076] in, This represents the row-level equivalent displacement; Indicates the amplitude; Indicates the frequency; Indicates the processing time; Indicates the processing line number; This indicates the exposure delay; Indicates the initial phase; Based on the row-level equivalent displacement and the vibration direction, the processed image frame is remapped to obtain a row-level corrected frame; the calculation formula is as follows:

[0077] in, This refers to the line-level correction frame; This refers to the processed image frame; Represents the x-coordinate of a pixel; Indicates the direction of the vibration; The fused image frame is obtained by weighted summation of the processed image frame and the row-level corrected frame using the region correction intensity; the calculation formula is as follows:

[0078] in, This refers to the fused image frame; The fused image frame is convolved with a Gaussian kernel, and a zero-constant is added as a normalized denominator. The corrected image is obtained by dividing the fused image frame by the normalized denominator. The calculation formula is as follows:

[0079] in, This refers to the corrected image; This refers to the Gaussian kernel; This represents the convolution operation; This represents the zero-resistance constant.

[0080] The total apparent magnification factor of candidate cracks within a region is the total number of apparent magnification factors corresponding to all candidate cracks extracted within a certain station component region, which is determined by the actual number of candidate cracks detected within that region.

[0081] The regional magnification mean is a dimensionless value obtained by arithmetically averaging the apparent magnification factors of all candidate cracks within a certain station component area. It characterizes the overall apparent pseudo-widening intensity of cracks in that area.

[0082] The region correction intensity is a dimensionless coefficient calculated from the region magnification mean. It is used to control the fusion weight between the original processed image frame and the row-level correction frame, and characterizes the degree of image correction in that region.

[0083] The pixel processing row number of the image frame is the sequence number of the pixel row used for row-level correction in the image frame. It represents the vertical position of the pixel row in the image frame and takes a positive integer value.

[0084] The row-level equivalent displacement is the equivalent imaging displacement of a pixel row in a certain image frame caused by the vibration of the station structure, and the unit is pixels.

[0085] The original processed image frame of a region is an uncorrected image frame captured by a camera device, corresponding to a specific station component region. It is the raw data for image restoration. It can be obtained by cropping the entire frame image captured by the camera device and extracting the corresponding pixel region image according to the predefined component region coordinates.

[0086] A row-level corrected frame is an image frame obtained by remapping the original processed image frame according to the row-level equivalent displacement and vibration direction. It is the basis of the corrected image after eliminating vibration artifacts.

[0087] The fused image frame is an image frame obtained by weighted summation of the original processed image frame and the row-level corrected frame using the region correction intensity, which takes into account both the original image details and the correction effect.

[0088] A Gaussian kernel with a scale of σ is a Gaussian convolution kernel used for image illumination normalization and is the core operator for achieving local grayscale smoothing in images. A two-dimensional Gaussian kernel with a kernel size of 5 x 5 pixels is preferred. This size is chosen because it effectively smooths local illumination fluctuations in the image without excessively blurring the edge features of cracks, thus adapting to the illumination characteristics of station structure images.

[0089] The scale parameter of the Gaussian kernel is the core parameter that determines the smoothness of the Gaussian kernel, representing the standard deviation of the Gaussian distribution. A value of 1.0 is preferred, as this value balances the effect of illumination normalization with the preservation of crack details, avoiding the loss of damage features due to excessive smoothing, and also preventing incomplete illumination normalization due to insufficient smoothing.

[0090] The second zero-prevention constant is a dimensionless, minimal constant added to the denominator to avoid zero values ​​in illuminance normalization calculations. It is preferably 0.0001, a value much smaller than the order of magnitude of the image grayscale values. Adding it will not affect the illuminance normalization calculation results, and it effectively avoids the problem of zero values ​​in the denominator.

[0091] The illuminance-normalized corrected image is an image frame obtained by performing Gaussian convolution and illuminance normalization on the fused image frame. It has a uniform illuminance distribution and eliminates vibration artifacts.

[0092] The convolution operator is a symbol that represents the convolution operation. The convolution operation is a pixel-by-pixel weighted summation calculation of the Gaussian kernel and the fused image frame, and it is the core operation for achieving local gray-level smoothing of an image.

[0093] In detail, the region correction intensity is calculated based on the average region magnification, giving regions with higher apparent pseudo-widening intensity a greater correction weight, thus achieving adaptive image restoration, rather than using a general fusion method with fixed weights. For example, if the average region magnification is 0.8, the region correction intensity is approximately 0.44, with line-level corrected frames accounting for 44% and original frames accounting for 56%, which can eliminate pseudo-widening while preserving original details. Based on the station structure vibration parameters and line exposure delay, line-level equivalent displacement is calculated, and line-level line-by-line remapping correction is performed on the progressively exposed images to accurately match the progressively exposed imaging characteristics of the camera equipment. Unlike traditional whole-frame image correction methods, this method effectively eliminates imaging distortion caused by line-by-line exposure coupling with structural vibration. It uses regional correction intensity to perform a weighted summation of the original processed image frame and the line-level correction frame to obtain a fused image frame, avoiding the loss of image details caused by directly replacing it with the correction frame, and adapting to the feature detection requirements of fine cracks in stations. After fusing the image frame with Gaussian kernel convolution, a second anti-zero constant is added as the normalization denominator to normalize the illumination, adapting to the actual scene of uneven illumination in stations. Unlike traditional global illumination normalization methods, this method can achieve local illumination uniformity and improve the accuracy of subsequent damage extraction.

[0094] In detail, a two-dimensional discrete Gaussian kernel is used, with a fixed kernel size of 5x5 pixels and a fixed scale parameter of 1.0. The convolution operation employs pixel-by-pixel two-dimensional convolution, and zero-padding is applied to image edges during calculation to avoid missing convolution calculations for edge pixels. Image frame-level remapping uses bilinear interpolation, which ensures the grayscale continuity of the remapped image and avoids pixelation distortion caused by nearest-neighbor interpolation, thus meeting the requirements for edge detection of fine cracks. The selection rule for the processed image frames is all image frames within the train's operating period, and these frames must be consistent with the train's operation. The frame sequences extracted during the line period remain consistent, requiring no additional filtering, ensuring the time matching between image restoration and train vibration; the second zero-prevention constant is fixed at 0.0001, requiring no adjustment based on component area or illumination conditions, and this value can adapt to the illumination normalization scenario of all station images; the weighted summation calculation of the fused image frames is performed pixel by pixel, and the gray value of each pixel is obtained by linearly weighting the gray values ​​of the corresponding pixels in the original processed image frame and the line-level correction frame according to the regional correction intensity, with the calculation accuracy retaining 6 decimal places to ensure the gray value accuracy of the fused image.

[0095] Preferably, the extraction of candidate cracks and peeling areas from the corrected image, and the measurement of the pixel width of the candidate cracks in the normal direction, includes: The horizontal and vertical discrete differences of the corrected image are calculated separately, and the square root of the sum of their squares is used to obtain the image gradient magnitude; the calculation formula is as follows:

[0096] in, This represents the gradient magnitude of the image; This refers to the corrected image; This represents the horizontal discrete difference; This represents the vertical discrete difference; and Represents pixel coordinates; The minimum eigenvalue of the second derivative matrix of the corrected image is calculated, and the larger of its negative value and zero is taken to obtain the linear edge response; the calculation formula is as follows:

[0097] in, This represents the linear edge response; Represents scale space set Extraction scale within; This represents the second derivative matrix corresponding to the extraction scale; Represents the minimum eigenvalue; The set of pixels whose image gradient magnitude is greater than the gradient determination threshold and whose linear edge response is greater than the response determination threshold is determined as the crack mask to extract the candidate cracks; the calculation formula is as follows:

[0098] in, This refers to the crack mask; Indicates characteristic functions; This represents the response determination threshold; This represents the gradient determination threshold; Calculate the local texture variance of the corrected image within the pixel neighborhood; the calculation formula is as follows:

[0099] in, This represents the local texture variance; Represents the variance operator; Represents the pixel neighborhood; The set of pixels whose local texture variance is greater than the variance determination threshold is identified as the peeling mask to extract the peeling area; the calculation formula is as follows:

[0100] in, This refers to the peeling mask; This represents the variance determination threshold; The crack mask is skeletonized to obtain a crack skeleton, and the pixel width of the candidate crack is measured along the normal direction of the crack skeleton.

[0101] Image gradient magnitude is the square root of the sum of the squares of the horizontal and vertical discrete differences at a pixel in the corrected image. It is dimensionless and characterizes the gradient intensity of the image edge at the pixel.

[0102] Horizontal discrete difference is the dimensionless change in grayscale value of a pixel along the horizontal direction in a corrected image.

[0103] Vertical discrete difference is the dimensionless change in grayscale value of a pixel along the vertical direction in a corrected image.

[0104] The crack linear edge response is the maximum positive value of the negative of the minimum eigenvalue of the second derivative matrix at different scales in the corrected image. It is dimensionless and characterizes the degree of crack edge response at the pixel that conforms to linear characteristics.

[0105] The scale space set for crack linear edge extraction is a set of different scale values ​​used to extract crack linear edges, and it is the core parameter controlling the scale of crack edge extraction. A scale value set of 1, 2, and 3 pixels is preferred, as this scale range can adapt to the different width characteristics of fine cracks in the station, achieving accurate extraction of fine crack linear edges without mis-extracting non-crack edges due to excessively large scales.

[0106] The second derivative matrix is ​​a matrix composed of the second-order partial derivatives of the corrected image at a specified scale.

[0107] The smallest eigenvalue is the eigenvalue with the smallest value obtained after eigenvalue decomposition of the second derivative matrix.

[0108] The crack linear edge response determination threshold is a dimensionless critical value used to filter crack linear edges. Only pixels with response values ​​greater than this threshold can be considered crack edges. Preferably, it is the 90th quantile of the crack linear edge response in the corrected image. This quantile effectively eliminates weak responses from background noise, preserves true crack linear edges, and adapts to the response characteristics of fine cracks in the station.

[0109] The image gradient magnitude threshold is a dimensionless critical value used to filter image edges; only pixels with gradient values ​​greater than this threshold can be considered image edges. Preferably, it is the 85th percentile of the corrected image gradient magnitude. This percentile can eliminate minor grayscale fluctuations in the background, retain true edge pixels, and reduce the false detection rate of crack extraction.

[0110] The crack mask is a set of binary pixels obtained by filtering the image gradient magnitude and the linear edge response of the crack with two thresholds. A pixel value of 1 represents a crack region and 0 represents a non-crack region. It is a binary image that represents the location of cracks in the image.

[0111] Local texture variance is the variance of gray values ​​of a pixel in a specified neighborhood of a corrected image. It is dimensionless and characterizes the uniformity of texture in the neighborhood of a pixel. The larger the variance, the more irregular the texture.

[0112] The pixel neighborhood centered on the coordinates is a square pixel region in the corrected image centered on a specified pixel, and it is the basic region for statistically analyzing the local texture variance. A 5x5 pixel square neighborhood is preferred, as this size balances the accuracy and computational efficiency of local texture statistics; it avoids distortion of statistical results due to an excessively small neighborhood, and prevents a surge in computational load due to an excessively large neighborhood.

[0113] The local texture variance threshold is a dimensionless critical value used to filter out peeling areas. Only pixels with a variance value greater than this threshold can be considered peeling areas. Preferably, it is the baseline mean of the corrected image's local texture variance plus 1.5 times the baseline standard deviation. This threshold effectively distinguishes between smooth areas on the structural surface and irregular texture areas caused by peeling, thus adapting to the texture characteristics of peeling on station structures.

[0114] The peeling mask is a set of binary pixels obtained by filtering local texture variance threshold. A pixel value of 1 represents the peeling area and 0 represents the non-peeling area. It is a binary image that represents the peeling location in the image.

[0115] The candidate crack pixel width along the crack skeleton normal is the pixel distance between sampling points on the crack skeleton along the normal. This can be obtained by skeletonizing the crack mask, drawing perpendicular lines along the normal at the skeleton sampling points, and counting the number of pixels between the intersections of the perpendicular lines and the crack mask edges.

[0116] In detail, damage extraction is first performed on the corrected image after restoration driven by the apparent magnification factor to eliminate imaging artifacts caused by the coupling of train vibration and line-by-line exposure. This separates the structural imaging problem from the damage semantic problem, avoiding the misjudgment of artifacts in the original image as damage features. For example, the corrected image of the station side wall has eliminated the edge blurring caused by vibration. At this time, extracting cracks can significantly reduce the false detection rate, which is different from the traditional method of directly extracting damage from the original image. A crack mask is extracted using a double threshold constraint of image gradient magnitude plus crack linear edge response. The gradient magnitude is used to filter edge pixels, and the linear edge response is used to filter linear pixels that meet the characteristics of fine cracks. Edge detection with dual constraints enables accurate extraction of fine cracks in stations, unlike general crack extraction methods that use a single edge detection operator. Local texture variance thresholding is employed to extract the peeling mask, leveraging the irregular texture and high variance of the peeling area to distinguish between peeling and smooth areas. This differs from traditional methods using grayscale thresholds and is adapted to the reality that peeling in station structures lacks fixed grayscale features. After skeletonizing the crack mask, pixel width is measured along the normal direction, transforming the planar region of the crack into a linear skeleton. Measurement along the normal direction accurately characterizes the actual width of the crack, unlike methods that directly count the mask pixel width and is adapted to the linear characteristics of fine cracks in stations.

[0117] In detail, the horizontal and vertical discrete differences of the corrected image are calculated using the Sobel operator, with a 3x3 pixel kernel used in both the horizontal and vertical directions. Zero-padding is applied to the image edges during calculation. The second derivative matrix is ​​calculated using the Hessian matrix method, obtained by calculating the second partial derivative of the image after Gaussian filtering. The scale of the filtered Gaussian kernel is consistent with the scale space set of the crack linear edge extraction. Morphological processing is required after crack and peeling mask extraction. First, a 1x1 pixel structuring element is used for opening operations to eliminate isolated noise points, and then a 2x2 pixel structuring element is used for closing operations to fill the tiny holes in the mask. The skeletonization of the crack mask uses a median transformation algorithm. The method extracts a linear region with a single pixel width, which can accurately characterize the direction of the crack. When measuring the pixel width along the normal of the crack skeleton, a perpendicular line with a length of 20 pixels is drawn along the normal at the skeleton sampling point. The number of pixels between the intersection of the perpendicular line and the edge of the crack mask is the pixel width at that point. Connectivity analysis is required for the crack and the peeling mask. Connectivity regions with fewer than 5 pixels are removed and identified as noise regions. The remaining connected regions are the final damaged regions. The calculation of each judgment threshold is based on the baseline data of the non-damaged corrected image. The baseline data is the statistical value of 100 frames of corrected images collected during the non-operation period of the train to ensure the objectivity and adaptability of the thresholds.

[0118] Preferably, the actual crack width and the actual peeling area are calculated by subtracting the artificially increased width from the pixel width based on the apparent magnification factor, including: The artificially increased width is obtained by multiplying the constant 4, the spatial scale, the amplitude, the absolute cosine of the angle difference between the vibration direction and the normal direction, the absolute sine of the product of pi, the frequency, and the camera exposure time, and the absolute sine of the product of pi, the frequency, the span value, and the exposure delay; the calculation formula is as follows:

[0119] in, This indicates the artificially increased width; Indicates the spatial scale; Indicates the amplitude; Indicates the direction of the vibration; Indicates the normal direction; Indicates the frequency; This indicates the camera exposure time; This indicates the value spanning multiple rows; This indicates the exposure delay; Multiply the spatial scale by the pixel width to obtain the initial physical width; subtract the artificially inflated width from the initial physical width to obtain the width difference; obtain a preset minimum width limit; compare the width difference with the minimum width limit, and select the larger value as the actual crack width; the calculation formula is as follows:

[0120] in, This indicates the actual crack width; This indicates the minimum width limit; Indicates the pixel width; The total number of peeled pixels within the peeled area is counted; the square of the spatial scale is calculated, and the square is multiplied by the total number of peeled pixels to obtain the actual peeled area; the calculation formula is as follows:

[0121] in, This represents the actual peeling area; This represents the total number of pixels that have been peeled off.

[0122] The artificially inflated width is caused by the coupling between the vibration of the station structure and the line-by-line exposure of the camera equipment. It is the artificially inflated value of the apparent width of the k-th candidate crack relative to the actual width, in millimeters.

[0123] The actual crack width is the true physical width of the crack obtained by subtracting the artificially inflated width at the m-th sampling point of the k-th candidate crack, in millimeters.

[0124] The preset minimum width limit is a minimum crack width set in millimeters to avoid negative crack widths after deducting the artificially inflated width. It is preferably 0.01 millimeters, which is the lower limit of accuracy for detecting fine cracks in stations. Adding this value avoids negative width issues in numerical calculations without substantially affecting the representation of the actual crack width, thus adapting to the detection scenario of fine cracks in stations.

[0125] The total number of peeled pixels in the peeled area is the first In the peeling mask of each peeling region, the total number of pixels representing the peeling region is obtained by counting the total number of pixels with a value of 1 in the peeling mask. The peeling mask is a binary image, where a pixel value of 1 represents a peeling region and 0 represents a non-peeling region.

[0126] The actual peeling area is the first The actual physical area of ​​the peeled region after deducting imaging artifacts is calculated, in square millimeters.

[0127] The initial physical width is obtained by directly multiplying the crack pixel width by the spatial scale, without deducting the artificially inflated crack physical width, and is in millimeters.

[0128] The width difference is the value obtained by subtracting the artificially increased width from the initial physical width. The unit is millimeters. It is the intermediate value for calculating the actual crack width, and the value may be positive or negative.

[0129] In detail, the dummy width is derived based on the core parameters of the apparent magnification factor specific to the station scene. This dummy width is precisely subtracted during the initial physical width conversion stage, rather than directly converting pixels to physical scale as in traditional methods. This fundamentally eliminates the systematic error in calibration conversion caused by the coupling of train vibration and line-by-line exposure. For example, if the initial physical width of a crack is 0.36 mm, the calculated dummy width is 0.115 mm. After subtraction, the actual crack width is 0.245 mm, restoring the true width of the crack. A preset minimum width is also introduced. The actual crack width is determined by comparing the width difference with this limit value and taking the larger value. This design is specifically for scenarios where the width difference of fine cracks in stations is prone to be negative, avoiding unreasonable results in numerical calculations. At the same time, this limit value is set as the lower limit of detection accuracy and does not affect the judgment of actual damage. The actual peeling area is calculated by multiplying the square of the spatial scale by the total number of peeling pixels. This not only adapts to the two-dimensional spatial characteristics of image pixels but also maintains consistency with the conversion scale of crack physical width, ensuring the uniformity of the physical quantity representation of station structural damage. This is different from the method of arbitrarily selecting the conversion ratio in the traditional peeling area calculation.

[0130] In detail, the preset minimum width limit is fixed at 0.01 mm, without adjustment based on station component type or crack characteristics. When the width difference is less than 0.01 mm, the output result is marked as less than 0.01 mm, while the value of 0.01 mm is uniformly used for subsequent risk calculations internally to avoid numerical instability caused by extremely small values. When counting the total number of peeling pixels in the peeling area, only regions with more than 5 connected pixels in the peeling mask are counted, and isolated noise points with less than or equal to 5 pixels are removed to ensure the accuracy of the statistical results. After the actual crack width is calculated, a sliding motion with 3 sampling points is used along the crack skeleton. The arithmetic mean is applied to the dynamic window to improve the smoothness and stability of the actual crack width value, avoiding the impact of calculation errors at a single sampling point on the overall judgment. When calculating the actual spalling area, adjacent connected regions in the spalling mask are merged. The adjacency criterion is that the minimum pixel distance between two connected regions is less than 3 pixels. The total area is calculated after merging, which is suitable for the characteristic of continuous distribution of spalling in station structures and ensures the integrity of the area statistics. The calculation accuracy of the actual crack width and the actual spalling area is retained to 3 decimal places, which not only meets the accuracy requirements of station structure damage detection, but also avoids the increase in calculation volume caused by excessive pursuit of accuracy.

[0131] Preferably, the actual crack width and the actual spalling area are summarized, and a score is calculated based on the preset importance of the load-bearing component to output a warning level, including: Extract the maximum value among all actual crack widths and determine it as the maximum crack width; the calculation formula is as follows:

[0132] in, This indicates the maximum crack width; Indicates the actual width of each of the aforementioned cracks; Sum all the actual peeling areas to obtain the cumulative peeling area; the calculation formula is as follows:

[0133] in, This represents the cumulative peeling area; Indicates the actual peeling area of ​​each item; The maximum crack width is divided by a preset width threshold and multiplied by a preset width weight to obtain a width risk item; the cumulative spalling area is divided by a preset area threshold and multiplied by a preset area weight to obtain an area risk item; the width risk item and the area risk item are added together to obtain a risk sum; the preset importance is multiplied by the risk sum to obtain the score; the calculation formula is as follows:

[0134] in, This indicates the score; This indicates the preset importance. This indicates the width weight; This indicates the width critical reference; This represents the area weight; This represents the critical area reference. The score is compared with both a first score threshold and a second score threshold. When the score is less than the first score threshold, a first warning level is output. When the score is greater than or equal to the first score threshold and less than the second score threshold, a second warning level is output. When the score is greater than or equal to the second score threshold, a third warning level is output. The determination formula is as follows:

[0135] in, The warning level refers to the level including the first warning level. Second warning level With the aforementioned third warning level ; This indicates the first score limit; This indicates the second score limit.

[0136] The maximum actual crack width within a region is the maximum value among all the actual crack widths at all sampling points within a certain station component region, expressed in millimeters, representing the most severe local crack damage within that region.

[0137] The cumulative actual spalling area within a region is the sum of the actual spalling areas of all spalling zones within a certain station component region, expressed in square millimeters, representing the overall degree of spalling damage development within that region.

[0138] The preset importance of load-bearing components is a dimensionless coefficient pre-set based on the load-bearing characteristics and structural sensitivity of the station's load-bearing components, used to characterize the damage risk weight of different load-bearing components. Preferably, 1.30 is used for the central column surface and column-slab joints, 1.20 for the corner openings and top beam-column joints, 1.00 for ordinary side wall strips, and 0.90 for the non-joint areas at the bottom of ordinary slabs. This value combines the characteristics of station structural engineering, with higher values ​​used for core load-bearing components and locally weak areas, so that the same damage will generate a higher risk score on important components, adapting to the damage control requirements of station structures.

[0139] The preset weight for crack width is a dimensionless weighting coefficient pre-defined for calculating the width risk item, representing the importance of crack width in damage risk assessment. A value of 0.50 is preferred because the width of fine cracks in station structure damage is a core indicator for determining structural safety; assigning a higher weight highlights the importance of crack width and aligns with the damage assessment principles for station structures.

[0140] The preset critical crack width benchmark is a pre-defined critical value used to normalize the maximum actual crack width, expressed in millimeters. It serves as the engineering benchmark for determining crack width risk. A value of 0.30 millimeters is preferred, as this is the engineering benchmark value for fine cracks in the concrete load-bearing components of the station that indicate the entry of them into a significant risk zone, meeting the requirements of the station structure safety inspection specifications.

[0141] The preset weight of the spalling area is a dimensionless weighting coefficient pre-set for calculating the area risk item, representing the importance of the spalling area in damage risk assessment. A value of 0.25 is preferred because the impact of spalling damage on the structural safety of the station is lower than that of crack width, maintaining consistency with the crack length weight to achieve a reasonable weight allocation for multiple damage indicators.

[0142] The preset critical benchmark for spalling area is a pre-defined threshold value used to normalize the cumulative actual spalling area, expressed in square millimeters. It serves as the engineering benchmark for determining the risk of spalling area. A preferred value is 2500 square millimeters, which is the engineering early warning benchmark value for single-area spalling damage to station concrete components, adapting to the spalling damage control requirements of station structures.

[0143] The structural damage risk score of a region is a dimensionless value calculated by combining the preset importance of the load-bearing component, the risk of crack width, and the risk of spalling area, which represents the overall risk level of structural damage in the region.

[0144] The first score threshold is a pre-set, dimensionless risk score threshold used to distinguish between the first and second warning levels. It is preferably 1, serving as the basic threshold for risk scores. Scores below this value indicate that the risk of damage is within a safe range, aligning with the basic grading logic of the three-level warning system.

[0145] The second score threshold is a pre-set, dimensionless risk score threshold used to distinguish between the second and third warning levels. It is preferably 2, which is the high-risk threshold for the risk score. A score higher than this value indicates that the damage risk is within the high-risk range, thus achieving a clear classification of the three warning levels.

[0146] The structural damage safety warning level of a region is determined by comparing the region's structural damage risk score with the score limit. It is used to intuitively characterize the degree of risk of structural damage to the station.

[0147] The width risk term is a dimensionless value obtained by multiplying the preset weight of the crack width by the ratio of the maximum actual crack width to the preset critical benchmark of the crack width, and it characterizes the specific damage risk caused by the crack width.

[0148] The area risk term is a dimensionless value obtained by multiplying the preset weight of the peeling area by the ratio of the cumulative actual peeling area to the preset critical benchmark for peeling area. It represents the specific damage risk caused by the peeling area.

[0149] The risk sum is a dimensionless value obtained by adding the width risk term and the area risk term, which represents the comprehensive risk level of cracks and spalling damage within the region.

[0150] In detail, a multi-indicator fusion method combining the maximum actual crack width and the cumulative actual spalling area is used to calculate damage risk, rather than the traditional single-indicator risk assessment method. This takes into account both the local severity of crack damage and the overall distribution characteristics of spalling damage. For example, if only crack damage exists in the central column area of ​​a station, its local risk can be accurately determined by the maximum actual crack width. The preset importance of load-bearing components is incorporated into the risk score calculation, and the load-bearing sensitivity and structural weakness of different load-bearing components in the station are explicitly encoded into the warning level. The central column, which bears the core load, and the corner openings that are locally weak are given higher importance coefficients, so that the same degree of damage presents different risk scores in different structural locations. Unlike traditional unified early warning and judgment methods without structural differences, this method normalizes the data by using the ratio of actual physical damage to the engineering critical benchmark. This achieves unified quantification of crack width and spalling area with different dimensions, allowing different types of damage risks to be directly added together. This differs from traditional non-normalized risk calculation methods. Furthermore, it employs a three-level early warning output with dual score limits, rather than a simple binary judgment of damage or no damage. This adapts to the graded management needs of station structural damage, ranging from safe to minor risk to high risk. For example, a risk score of 0.7978 is classified as Level 1, and a score of 1.0465 as Level 2, providing station maintenance personnel with clear and detailed handling guidelines.

[0151] In detail, the preset importance of load-bearing components is fixed according to the type of station components: 1.30 for central column surfaces and column-slab nodes, 1.20 for corner openings and top-floor beam-column nodes, 1.00 for ordinary sidewalls, and 0.90 for non-node areas at the bottom of ordinary slabs. This value covers all major concrete load-bearing components of the station and does not require dynamic adjustment based on damage conditions. The preset weight for crack width is 0.50, and the preset weight for spalling area is 0.25, with an additional 0.25 weight reserved for cumulative crack length, achieving a three-dimensional damage risk assessment of crack width, crack length, and spalling area. The preset critical benchmarks for crack width are 0.30 mm, for spalling area 2500 square millimeters, and for cumulative crack length 1000 mm, all of which are built-in engineering benchmark values ​​for station structural damage and do not require on-site debugging. The first score threshold is fixed at 1, and the second score threshold is fixed at 2. The corresponding warning levels are Level 1 (Safe), Level 2 (Caution), and Level 3 (Warning). When outputting the warning level, the specific risk score and core damage indicators are simultaneously marked. When calculating the cumulative actual peeling area, adjacent peeling areas with a pixel spacing of less than 5 within the same component area are merged before calculating the total area to avoid duplicate counting of scattered peeling areas and ensure the completeness of the area statistics. The risk score calculation accuracy is retained to 4 decimal places. When determining the warning level, the actual score value is directly compared with the threshold without rounding. The output frequency of the warning level is consistent with the detection frequency during train operation. The warning level of the detection is output immediately after the train passes through the station. During non-train operation periods, the damage warning level of the component area is output every 30 minutes to ensure the real-time operation and maintenance of the station.

[0152] Example 2: A station structure damage identification method based on image recognition, applied to any of the image recognition-based station structure damage identification systems described above, including: The image sequence of the carrier is acquired by the camera equipment, the train operation period is extracted by the gray scale change of the preset reference area, and the pixel scale is converted into the spatial scale. During the train's operation, the imaging displacement is obtained by tracking the non-damaged parts of the reference area, and the vibration direction, amplitude, frequency, and initial phase are determined. Candidate cracks are obtained from the image sequence, and the normal direction and span value of the candidate cracks are extracted. The apparent magnification factor is calculated by combining the spatial scale, vibration direction, amplitude, frequency, initial phase and exposure delay of the camera device. The image sequence is then corrected using the apparent magnification factor to obtain a corrected image. The candidate cracks and peeling areas are extracted from the corrected image, and the pixel width of the candidate cracks is measured in the normal direction. Based on the apparent magnification factor, the artificially increased width is subtracted from the pixel width to calculate the actual crack width and the actual peeling area. The actual crack width and the actual spalling area are combined, and a score is calculated based on the preset importance of the load-bearing component, and an early warning level is output.

[0153] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0154] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.

Claims

1. A station structure damage identification system based on image recognition, characterized in that, include: The acquisition and extraction module acquires image sequences of the carrier through camera equipment, extracts the train operation period using grayscale changes in a preset reference area, and converts the pixel scale into a spatial scale. The vibration estimation module obtains the imaging displacement by tracking the non-damaged parts of the reference area during the train's operation period, and calculates the vibration direction, amplitude, frequency, and initial phase. The restoration module obtains candidate cracks from the image sequence, extracts the normal direction and span value of the candidate cracks, and calculates the apparent magnification factor by combining the spatial scale, vibration direction, amplitude, frequency, initial phase and exposure delay of the camera device. The image sequence is then corrected using the apparent magnification factor to obtain a corrected image. The detection module extracts the candidate cracks and peeling areas from the corrected image and measures the pixel width of the candidate cracks in the normal direction. The compensation module calculates the actual crack width and the actual peeling area by subtracting the artificially increased width from the pixel width based on the apparent magnification factor. The early warning module summarizes the actual crack width and the actual spalling area, calculates a score based on the preset importance of the load-bearing component, and outputs the early warning level.

2. The station structure damage identification system based on image recognition according to claim 1, characterized in that, Image sequences of the carrier components are acquired using camera equipment. The train's operating period is extracted using grayscale changes in a preset reference area, and the pixel scale is converted into a spatial scale, including: Obtain the current frame image and the previous frame image from the image sequence, and calculate the absolute value of the grayscale difference between all pixels in the reference area and the current frame image and the previous frame image; Calculate the average of the absolute values ​​of all the grayscale differences to obtain the average inter-frame energy. Obtain the mean baseline energy and standard deviation of the reference area under historical stable conditions, and add the mean baseline energy to twice the standard deviation of the baseline energy to obtain the operational period judgment boundary; The set of consecutive image frames whose average inter-frame energy is greater than the operating period determination threshold is determined as the train operating period; Obtain the known reference length within the reference region and the image pixel span corresponding to the known reference length; Divide the known reference length by the image pixel span to obtain the converted spatial scale.

3. The station structure damage identification system based on image recognition according to claim 2, characterized in that, During the train's operation, imaging displacement is obtained by tracking the non-damaged portion of the reference area, and the vibration direction, amplitude, frequency, and initial phase are determined, including: Multiple reference tracking points are set at the non-damaged portion of the reference region, and the two-dimensional imaging displacement vector of each reference tracking point is obtained. Calculate the time mean displacement vector of each of the reference tracking points and the overall mean displacement vector of all the reference tracking points, and then calculate the displacement covariance matrix. Extract the eigenvector corresponding to the largest eigenvalue of the displacement covariance matrix, and determine the direction of the eigenvector as the vibration direction; The two-dimensional imaging displacement vector of each of the reference tracking points is projected onto the vibration direction to obtain the projected displacement. By fitting the projected displacement of each of the reference tracking points with a sine function, the single-point vibration amplitude, single-point vibration frequency and single-point initial phase of each of the reference tracking points are obtained. Based on the preset tracking point weights, the single-point vibration amplitude, single-point vibration frequency, and single-point initial phase of each reference tracking point are calculated by weighted average to obtain the amplitude, frequency, and initial phase, respectively.

4. The station structure damage identification system based on image recognition according to claim 3, characterized in that, Extract the normal direction and span value of the candidate cracks, and combine them with the spatial scale, vibration direction, amplitude, frequency, initial phase, and exposure delay of the camera device to calculate the apparent magnification factor, including: Edge extraction is performed on the initial image frames in the image sequence to obtain the candidate cracks, and the number of sampling points of the candidate cracks and the normal pixel width of each sampling point are obtained. The average width of all the normal pixels is multiplied by the spatial scale to obtain the average width of the candidate cracks; Obtain the video exposure time of the camera device, and calculate the absolute cosine value of the angle difference between the vibration direction and the normal direction; Calculate the absolute value of the sine of the product of pi, the frequency, and the camera exposure time, and use it as the first absolute value of the sine. Calculate the absolute sine of the product of pi, the frequency, the span value, and the exposure delay, and use it as the second absolute sine. Multiply the constant 4, the spatial scale, the amplitude, the absolute cosine of the angle difference, the absolute first sine, and the absolute second sine to obtain the amplification numerator; Add the average width of the candidate cracks to the first zero-prevention constant to obtain the amplified denominator term; The apparent amplification factor is obtained by dividing the amplification numerator by the amplification denominator.

5. The station structure damage identification system based on image recognition according to claim 4, characterized in that, Correcting the image sequence using the apparent magnification factor to obtain a corrected image includes: The average value of all the apparent magnification factors is calculated as the regional magnification mean. The regional correction intensity is obtained by dividing the regional magnification mean by a constant and the sum of the regional magnification mean. Calculate the product of the exposure delay and the processing row number of the processed image frame in the image sequence, and add it to the processing time to obtain the time variable; Multiply the constant 2, pi, the frequency, and the time variable together, add the result to the initial phase, take the sine value, and then multiply it by the amplitude to obtain the row-level equivalent displacement. Based on the row-level equivalent displacement and the vibration direction, the processed image frame is remapped to obtain a row-level corrected frame; The processed image frame and the row-level corrected frame are weighted and summed using the region correction intensity to obtain a fused image frame; The fused image frame is convolved with a Gaussian kernel and a second zero constant is added as a normalized denominator. The fused image frame is then divided by the normalized denominator to obtain the corrected image.

6. The station structure damage identification system based on image recognition according to claim 5, characterized in that, The candidate crack and peeling area are extracted from the corrected image, and the pixel width of the candidate crack is measured in the normal direction, including: The horizontal and vertical discrete differences of the corrected image are calculated separately, and the square root of the sum of their squares is used to obtain the image gradient magnitude. Calculate the minimum eigenvalue of the second derivative matrix of the corrected image, and take the larger of its negative value and zero to obtain the linear edge response; The set of pixels whose image gradient magnitude is greater than the gradient determination threshold and whose linear edge response is greater than the response determination threshold is determined as the crack mask to extract the candidate crack; Calculate the local texture variance of the corrected image in the pixel neighborhood; The set of pixels whose local texture variance is greater than the variance determination threshold is identified as the peeling mask to extract the peeling area; The crack mask is skeletonized to obtain a crack skeleton, and the pixel width of the candidate crack is measured along the normal direction of the crack skeleton.

7. The station structure damage identification system based on image recognition according to claim 6, characterized in that, Based on the apparent magnification factor, the artificially increased width is subtracted from the pixel width to calculate the actual crack width and the actual peeling area, including: The artificially increased width is obtained by multiplying the constant 4, the spatial scale, the amplitude, the absolute cosine of the angle difference, the absolute first sine, and the absolute second sine. Multiply the spatial scale by the pixel width to obtain the initial physical width; Subtract the artificially increased width from the initial physical width to obtain the width difference; Obtain a preset minimum width limit, compare the width difference with the minimum width limit, and select the larger value to determine the actual crack width; Count the total number of peeled pixels within the peeled area; Calculate the square of the spatial scale, multiply the square by the total number of peeled pixels, and obtain the actual peeled area.

8. The station structure damage identification system based on image recognition according to claim 7, characterized in that, The actual crack width and the actual spalling area are summarized, and a score is calculated based on the preset importance of the load-bearing component to output a warning level, including: Extract the maximum value among all the actual crack widths and determine it as the maximum crack width; Sum all the actual peeling areas to obtain the cumulative peeling area; Divide the maximum crack width by a preset width critical benchmark and multiply it by a preset width weight to obtain the width risk term; Divide the cumulative peeling area by a preset area threshold and multiply it by a preset area weight to obtain the area risk item. The sum of the width risk item and the area risk item is added together to obtain the risk value. The score is obtained by multiplying the preset importance of the bearing component by the risk value. The scores are compared with the first score limit and the second score limit, respectively; When the score is less than the first score limit, the warning level is determined to be the first warning level; When the score is greater than or equal to the first score limit and less than the second score limit, the warning level is determined to be the second warning level; When the score is greater than or equal to the second score limit, the warning level is determined to be the third warning level.

9. A station structure damage identification method based on image recognition, applied to the station structure damage identification system based on image recognition as described in any one of claims 1-8, characterized in that, include: The image sequence of the carrier is acquired by the camera equipment, the train operation period is extracted by the gray scale change of the preset reference area, and the pixel scale is converted into the spatial scale. During the train's operation, the imaging displacement is obtained by tracking the non-damaged parts of the reference area, and the vibration direction, amplitude, frequency, and initial phase are determined. Candidate cracks are obtained from the image sequence, and the normal direction and span value of the candidate cracks are extracted. The apparent magnification factor is calculated by combining the spatial scale, vibration direction, amplitude, frequency, initial phase and exposure delay of the camera device. The image sequence is then corrected using the apparent magnification factor to obtain a corrected image. The candidate cracks and peeling areas are extracted from the corrected image, and the pixel width of the candidate cracks is measured in the normal direction. Based on the apparent magnification factor, the artificially increased width is subtracted from the pixel width to calculate the actual crack width and the actual peeling area. The actual crack width and the actual spalling area are combined, and a score is calculated based on the preset importance of the load-bearing component, and an early warning level is output.