Catenary component thermal defect detection method based on space-time fusion and speed normalization

By employing spatiotemporal fusion and vehicle speed normalization, the problems of vehicle speed fluctuation and inconsistency of multi-source data in the thermal defect detection of overhead contact system components are solved, achieving high-precision, low-false-alarm thermal defect identification, which is suitable for stable detection in complex environments.

CN121456768BActive Publication Date: 2026-04-14CHENGDU NAT RAILWAYS ELECTRICAL EQUIP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies for detecting thermal defects in overhead contact system components suffer from problems such as inconsistent data due to vehicle speed fluctuations, inconsistent timestamps from multiple data sources, inaccurate feature judgment of single-frame images, lack of utilization of time sequence information, and misjudgment due to environmental interference, resulting in unstable detection results and a high false alarm rate.

Method used

By employing spatiotemporal fusion and vehicle speed normalization methods, image sequences and high-frequency speed data are acquired using vehicle-mounted infrared image acquisition equipment and Doppler odometers. High-precision position sequences are generated by combining time synchronization protocols, cross-frame correlation is performed using a target detection model, spatiotemporal features are extracted, and consistency verification is conducted through a decision model to output thermal defect judgment results.

Benefits of technology

It significantly improves the stability and accuracy of detection, reduces misjudgments due to environmental interference, adapts to detection under different vehicle speed conditions, and ensures the reliability and consistency of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456768B_ABST
    Figure CN121456768B_ABST
Patent Text Reader

Abstract

The application discloses a kind of contact net component thermal defect detection methods based on space-time fusion and speed normalization, belong to railway electrification detection and intelligent perception field.The method first acquires contact net component image sequence, high-frequency speed and low-frequency position and speed information, by space-time synchronization unified time stamp and generates high-precision position sequence, realizes standardization in the same physical distance acquisition frame by speed normalization;Again, target detection and cross-frame tracking are carried out on the normalized image to form a time sequence trajectory, the spatial and temporal features of the trajectory are extracted and fused, and the decision model is input combined with the space-time consistency constraint to judge the thermal defects.The method effectively solves the problems of inconsistent time sequence features, poor multi-source data matching, and susceptibility to interference and misjudgment in traditional detection, significantly improves the stability, accuracy and anti-interference ability of thermal defect detection, and adapts to dynamic operating conditions of trains.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of railway electrification inspection and intelligent sensing, and in particular to a method for detecting thermal defects in catenary components based on spatiotemporal fusion and vehicle speed normalization. Background Technology

[0002] With the rapid development of the rail transit industry towards intelligent and unmanned operation and maintenance, the overhead contact system, as a core component of the train power supply system, has become a crucial link in ensuring line safety through real-time monitoring and fault early warning of its operating status. Currently, the industry widely adopts onboard inspection solutions for overhead contact system components. The core technology is based on infrared imaging, which continuously acquires temperature distribution images of overhead contact system components, including disconnect switches, insulators, and dropper clamps, using onboard infrared equipment. Temperature anomalies are the core basis for judging thermal defects. Simultaneously, auxiliary data such as train speed and location along the line are acquired in conjunction with onboard speed detection devices and train management systems (TMS), forming a multi-source data acquisition architecture of "image + sensor." Furthermore, to improve the level of intelligent inspection, the industry is gradually introducing technologies such as deep learning target detection and multi-source data fusion to try to solve the problems of component positioning and status recognition in complex scenarios. Infrared equipment also needs to adapt to dynamic conditions such as train bumps and turns to ensure stable lens coverage of key areas of the overhead contact system. The overall technical direction is continuously optimized around the inspection requirements of "more accurate, more real-time, and more reliable."

[0003] Current technologies for detecting thermal defects in overhead contact system components still face several challenges requiring breakthroughs: During train operation, significant speed fluctuations lead to an excessive number of frames for the same component in infrared image sequences at low speeds, increasing data storage and processing pressure and hindering temporal feature extraction due to redundant information. At high speeds, insufficient frames result in missing crucial temperature information, making it impossible to provide standardized input data for subsequent time-series analysis. Furthermore, multi-source data, including infrared images, speed data, and location information, lacks a unified spatiotemporal synchronization mechanism. Inconsistent timestamps across different devices can cause misalignments between a single image frame and the corresponding speed and location data, hindering effective collaborative correction of detection results. Finally, detection logic generally relies on single-frame images. The spatial temperature characteristics used to determine thermal defects rely solely on parameters such as the average temperature within a single frame and the proportion of high-temperature areas. This approach fails to fully utilize temporal information such as temperature change trends and component movement trajectories across consecutive frames, making it difficult to distinguish between temporary thermal interference such as solar reflection and instantaneous noise and genuine thermal defects. Furthermore, the thermal defect decision-making process lacks a systematic constraint mechanism. It fails to perform dual verification on the spatial distribution concentration of high-temperature areas, such as whether they are concentrated in critical conductive parts of components, and the temporal stability of temperature changes, such as whether they are continuously in a high-temperature state. This leads to some interference signals being misjudged as thermal defects, affecting the reliability of the detection results and making it difficult to meet the actual operation and maintenance requirements for low false alarms and high accuracy. Especially in complex environments, the fluctuation of detection results is further amplified, failing to provide stable support for operation and maintenance decisions. Summary of the Invention

[0004] The purpose of this invention is to overcome the problems of the prior art and provide a method for detecting thermal defects in catenary components based on spatiotemporal fusion and vehicle speed normalization.

[0005] The objective of this invention is achieved through the following technical solution:

[0006] A method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization is provided. The method includes the following steps:

[0007] S1. Continuously acquire image sequences of the overhead contact line components using on-board acquisition equipment, and simultaneously acquire high-frequency speed data provided by the on-board speed detection device and low-frequency position and speed information provided by the train position and speed device;

[0008] S2. Based on the time synchronization protocol, the timestamps of the image sequence, high-frequency speed data, low-frequency position information and low-frequency speed information are unified. The low-frequency position information and high-frequency speed data are combined to generate a position sequence corresponding to the image sequence. Based on the vehicle speed information, the image sequences at different vehicle speeds are mapped to an equidistant sampling spatial scale to achieve normalization of the acquired frames within the same physical distance.

[0009] S3. A target detection model is used to detect the catenary components in the normalized image sequence. Combining the spatial location and motion prediction information of the detection results, a tracking algorithm is used to achieve cross-frame target association and form a set of temporal trajectories of the same catenary component.

[0010] S4. For each time-series trajectory, extract the spatial and temporal features of the target area, and fuse the spatial and temporal features to form a spatiotemporal fusion feature;

[0011] S5. Input the spatiotemporal fusion features into the decision model, combine them with the spatiotemporal consistency constraints to determine whether there are thermal defects in the contact wire components, and output the thermal defect judgment results.

[0012] Furthermore, in S1, the on-board acquisition device is an infrared image acquisition device, and the image sequence is an infrared image sequence; the on-board speed detection device is a Doppler odometer, and the train position and speed device is a train TMS system.

[0013] Furthermore, in step S2, the low-frequency position information and high-frequency velocity data are combined to generate a position sequence corresponding to the image sequence. Specifically, this includes: using the low-frequency position information as a reference anchor point and the high-frequency velocity data as an interpolation basis, the low-frequency position information is used as a position reference. The velocity change between two adjacent low-frequency position information points is interpolated and supplemented by the high-frequency velocity data. Then, the interpolated velocity signal is integrated to obtain high-precision position data corresponding to each frame of the image sequence, thereby forming the position sequence.

[0014] Furthermore, in step S2, the image sequences at different vehicle speeds are mapped to an equidistant sampling spatial scale based on the vehicle speed information. Specifically, this includes: presetting a target vehicle speed and a target number of frames; if the current vehicle speed is lower than the target vehicle speed and the original number of frames in the image sequence is greater than the target number of frames, then an equidistant frame extraction method is used to extract images of the target number of frames from the original image sequence; if the current vehicle speed is higher than the target vehicle speed and the original number of frames in the image sequence is less than the target number of frames, then an equidistant frame interpolation method is used to insert images into the original image sequence to make up for the target number of frames; if the current vehicle speed is equal to the target vehicle speed, then the original image sequence is used directly.

[0015] Furthermore, in S3, the target detection model is a deep learning-based target detection model, and the contact wire component includes at least one of disconnecting switches, insulators, and dropper clamps.

[0016] Furthermore, in S3, the motion prediction information is calculated through a motion model, which is a Kalman filter model; the correlation basis of the tracking algorithm includes at least one of the following: the motion distance between the detection result and the predicted position, the similarity of appearance features between the detected target and the tracked target, and the spatial overlap between the detection box and the predicted box.

[0017] Furthermore, in S4, the spatial features include at least one of the following: the average temperature of the target area, the temperature variance, the temperature gradient distribution, the contrast between the target area and the neighboring pixel areas, and the area ratio of the high-temperature area to the target area.

[0018] Furthermore, in S4, the temporal features include at least one of the following: the temperature change curve of the target region in consecutive frames, the temperature change slope, the temperature stability index, and the temporal change trend of the area ratio of high-temperature regions.

[0019] Furthermore, in S5, the decision model is an anomaly detection model or a classification model; the spatiotemporal consistency constraint includes spatial consistency constraint and temporal consistency constraint, wherein the spatial consistency constraint is the concentration of high temperature region distribution and the stability of target-background contrast, and the temporal consistency constraint is the stability of temperature change trend and the degree of temperature fluctuation.

[0020] Furthermore, in step S4, when fusing the spatial and temporal features, environmental context features are also incorporated; the environmental context features are obtained by combining the location sequence with a preset scene database, including at least one of the following: the environmental type of the contact wire component, the ambient background radiation value, and the ambient temperature.

[0021] The beneficial effects of this invention are:

[0022] (1) By implementing spatiotemporal synchronization, vehicle speed normalization, target cross-frame tracking, spatiotemporal feature fusion and spatiotemporal consistency constraints in a coordinated manner, the real thermal defects and instantaneous interference can be effectively distinguished, the problem of inconsistent timing features in traditional detection can be solved, and the stability and continuity of detection can be significantly improved.

[0023] (2) By extracting the spatial and temporal features of the target area and integrating environmental context information, combined with dual consistency verification, the misjudgment caused by environmental interference is reduced, and the accuracy and reliability of thermal defect identification are greatly improved.

[0024] (3) By unifying the time of multi-source data and generating high-precision location sequences, the standardized adaptation of data under different operating conditions can be achieved, the scope of application of detection scenarios can be broadened, and the detection process can be promoted efficiently and smoothly. Attached Figure Description

[0025] Figure 1 A flowchart of a method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization;

[0026] Figure 2 Example of false alarm due to solar reflection - temperature change curve;

[0027] Figure 3Example of false solar reflection - curve of bright spot rate variation;

[0028] Figure 4 A schematic diagram illustrating the detection effect of the contact wire component provided in the embodiment;

[0029] Figure 5 The main flowchart for thermal defect detection provided in this embodiment;

[0030] Figure 6 The flowchart for vehicle speed normalization processing provided in the embodiment;

[0031] Figure 7 A flowchart of target detection and tracking provided for an embodiment;

[0032] Figure 8 The spatiotemporal feature fusion and decision-making flowchart is provided for the embodiment. Detailed Implementation

[0033] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] Example 1

[0035] See Figure 1 A method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization is provided, including the following steps:

[0036] S1. Continuously acquire image sequences of the overhead contact line components using on-board acquisition equipment, and simultaneously acquire high-frequency speed data provided by the on-board speed detection device and low-frequency position and speed information provided by the train position and speed device;

[0037] S2. Based on the time synchronization protocol, the timestamps of the image sequence, high-frequency speed data, low-frequency position information and low-frequency speed information are unified. The low-frequency position information and high-frequency speed data are combined to generate a position sequence corresponding to the image sequence. Based on the vehicle speed information, the image sequences at different vehicle speeds are mapped to an equidistant sampling spatial scale to achieve normalization of the acquired frames within the same physical distance.

[0038] S3. A target detection model is used to detect the catenary components in the normalized image sequence. Combining the spatial location and motion prediction information of the detection results, a tracking algorithm is used to achieve cross-frame target association and form a set of temporal trajectories of the same catenary component.

[0039] S4. For each time-series trajectory, extract the spatial and temporal features of the target area, and fuse the spatial and temporal features to form a spatiotemporal fusion feature;

[0040] S5. Input the spatiotemporal fusion features into the decision model, combine them with the spatiotemporal consistency constraints to determine whether there are thermal defects in the contact wire components, and output the thermal defect judgment results.

[0041] In step S1, the on-board acquisition device is an infrared image acquisition device, and the image sequence is an infrared image sequence; the on-board speed detection device is a Doppler odometer, and the train position and speed device is a train TMS system.

[0042] Step S2 combines the low-frequency position information and high-frequency velocity data to generate a position sequence corresponding to the image sequence. Specifically, it includes: using the low-frequency position information as a reference anchor point and the high-frequency velocity data as an interpolation basis, using the low-frequency position information as a position reference, interpolating and supplementing the velocity change between two adjacent low-frequency position information points using high-frequency velocity data, and then performing integration on the interpolated velocity signal to obtain high-precision position data corresponding to each frame of the image sequence, thereby forming the position sequence.

[0043] In step S2, the image sequences at different vehicle speeds are mapped to an equidistant sampling spatial scale based on the vehicle speed information. Specifically, this includes: setting a target vehicle speed and a target number of frames; if the current vehicle speed is lower than the target vehicle speed and the original number of frames in the image sequence is greater than the target number of frames, an equidistant frame extraction method is used to extract images of the target number of frames from the original image sequence; if the current vehicle speed is higher than the target vehicle speed and the original number of frames in the image sequence is less than the target number of frames, an equidistant frame interpolation method is used to insert images into the original image sequence to make up for the target number of frames; if the current vehicle speed is equal to the target vehicle speed, the original image sequence is used directly.

[0044] In step S3, the target detection model is a deep learning-based target detection model, and the contact wire component includes at least one of disconnecting switches, insulators, and dropper clamps.

[0045] In step S3, the motion prediction information is calculated by a motion model, which is a Kalman filter model; the correlation basis of the tracking algorithm includes at least one of the following: the motion distance between the detection result and the predicted position, the similarity of appearance features between the detected target and the tracked target, and the spatial overlap between the detection box and the predicted box.

[0046] In step S4, the spatial features include at least one of the following: the average temperature of the target area, the temperature variance, the temperature gradient distribution, the contrast between the target area and the neighboring pixel areas, and the area ratio of the high-temperature area to the target area.

[0047] In step S4, the temporal features include at least one of the following: the temperature change curve of the target region in consecutive frames, the temperature change slope, the temperature stability index, and the temporal change trend of the area ratio of high-temperature regions.

[0048] In step S5, the decision model is an anomaly detection model or a classification model; the spatiotemporal consistency constraint includes spatial consistency constraint and temporal consistency constraint, wherein the spatial consistency constraint is the concentration of high temperature region distribution and the stability of target-background contrast, and the temporal consistency constraint is the stability of temperature change trend and the degree of temperature fluctuation.

[0049] In step S4, when fusing the spatial and temporal features, environmental context features are also incorporated; the environmental context features are obtained by combining the location sequence with a preset scene database, including at least one of the following: the environmental type of the contact wire component, the ambient background radiation value, and the ambient temperature.

[0050] Example 2

[0051] A specific embodiment of a method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization is provided:

[0052] Traditional vehicle-mounted infrared detection systems suffer from numerous limitations in detecting thermal defects in overhead contact line components: they rely solely on single-frame infrared images or single-point temperature values ​​for judgment, resulting in a limited detection dimension and susceptibility to environmental interference such as sunlight reflection, leading to false positives; they fail to fully utilize temporal information, with significant differences in the number of frames acquired for the same component at different vehicle speeds, resulting in inconsistent temporal feature lengths and hindering stable detection; the low update frequency of the TMS broadcast signal cannot meet the spatiotemporal synchronization requirements of high-frame-rate infrared video, affecting the accurate matching of multi-source data; and they lack intelligent recognition and compensation mechanisms for environmental scenes, leading to significant performance variations under different environments. These problems result in high false alarm rates and poor detection stability using traditional methods, necessitating improvements in detection accuracy through multi-dimensional feature fusion and temporal consistency verification.

[0053] like Figure 2 As shown, in traditional single-frame detection methods, areas of sunlight reflection in infrared images (marked with boxes) are misjudged as thermal defects due to abnormally high local temperatures; however, this embodiment uses temporal consistency analysis (… Figure 2 (b) shows the temperature change curve. It was found that the temperature in this area fluctuated drastically over time and eventually returned to normal levels, thus accurately identifying it as a transient disturbance and eliminating it. Figure 3 As shown, traditional single-frame detection also misidentifies solar reflective areas (marked with boxes) as thermal defects. This embodiment utilizes the temporal characteristics of bright spot rate (…). Figure 3(b) shows the bright spot rate curve. The bright spot rate refers to the proportion of high-temperature areas in the infrared image to the total area of ​​the component. The method is used for discrimination, and the results show that the proportion of high-temperature areas in this region changes significantly and has poor persistence, thus achieving accurate identification and elimination of interfering targets. The specific implementation of this method is as follows:

[0054] Step 1: Data Collection and Acquisition of Multi-Source Information

[0055] To achieve accurate detection of thermal defects in overhead contact system components, it is first necessary to simultaneously collect multi-source data. This provides a foundation for subsequent spatiotemporal synchronization, vehicle speed normalization, and feature analysis. (See [reference needed]) Figure 5 This diagram fully illustrates the end-to-end processing logic from data input to decision output. This section represents the "data acquisition" step, and the specific implementation process is as follows:

[0056] Infrared images of the overhead contact system components are continuously captured at a frame rate of 70Hz using an onboard infrared camera. Each frame is embedded with a precise microsecond-level timestamp, which serves as the benchmark for all subsequent data synchronization. During the acquisition process, the infrared camera lens is pointed towards the overhead contact system area at a fixed angle to the direction of train operation. This ensures stable imaging of critical components such as disconnect switches, upper and lower anchor section insulators, and dropper clamps, and that the images fully reflect the temperature distribution characteristics of the component surfaces—normal components have temperatures close to ambient temperatures, while components with thermal defects will exhibit localized high temperatures in critical areas such as conductive connections.

[0057] The system simultaneously acquires high-frequency continuous speed data (update frequency 100Hz) from the onboard Doppler odometer and low-frequency absolute position and speed information (update frequency 1Hz) from the train's TMS system. Both data carry precise time stamps. The Doppler odometer calculates real-time speed by sensing changes in the rotational speed of the train's wheel axles and combining this with preset wheel diameter parameters. Its data update frequency is much higher than that of the TMS system, enabling it to accurately capture instantaneous speed fluctuations during acceleration, deceleration, and constant speed phases, providing high-density data support for speed interpolation in subsequent speed normalization processing. The absolute position information output by the TMS system (such as mileage markers along the line) has absolute reference value and can serve as an "anchor point" for subsequent "high-precision position sequence generation," correcting potential position accumulation errors during high-frequency speed data integration. Its low-frequency speed information can also serve as auxiliary verification of the Doppler odometer's high-frequency speed data. If the speed deviation between the two systems exceeds a preset threshold within the same time interval, a speed data calibration mechanism can be triggered, using the average speed output by the TMS system as a benchmark for linear correction of the Doppler speed data.

[0058] Step 2: Spatiotemporal synchronization and vehicle speed normalization processing:

[0059] To eliminate the impact of time skew and vehicle speed variations on detection from multi-source data, spatiotemporal synchronization and vehicle speed normalization are required to convert non-standardized data into input with a uniform scale. (See [reference needed]). Figure 6 The diagram shows the specific process for vehicle speed normalization. The detailed implementation process is as follows:

[0060] Based on a time synchronization protocol, the timestamps of infrared image sequences, Doppler odometer high-frequency velocity data, and TMS system low-frequency position and velocity information are uniformly calibrated. First, the acquisition units of each data source record the timestamp of the original data generation to ensure the accuracy of time recording—the timestamp of the infrared image is stored in the image metadata, while the timestamps of the Doppler velocity data and TMS information are stored in the header of the corresponding data frames. Then, using the TMS system time as the global reference clock (synchronized with the train dispatch center with an error of less than 1ms), the timestamps of the infrared cameras and Doppler odometers are adjusted through the time synchronization protocol, mapping all data to the same time coordinate system. After calibration, multiple sets of data are randomly selected to verify time consistency, ensuring that each frame of infrared image, each velocity data point, and each piece of position information corresponds to a unique and consistent time node.

[0061] Using low-frequency position information provided by the TMS system as reference anchor points and high-frequency velocity data provided by the Doppler odometry as interpolation basis, a high-precision position sequence corresponding one-to-one with the infrared image sequence is generated. The specific steps are as follows: extract "time-position" anchor point pairs from the TMS data; within the time interval corresponding to two adjacent anchor points, interpolation is performed using high-frequency velocity data; if the velocity changes linearly, linear interpolation is used; if it changes non-linearly, cubic spline interpolation is used to ensure that the interpolated velocity curve smoothly reflects the actual velocity change; the interpolated continuous velocity signal is integrated at time intervals (matching the infrared image frame rate, i.e., 1 / 70 ≈ 0.0143 s) to calculate the position offset of each time node relative to the previous anchor point; combined with the absolute position of the previous anchor point, the absolute position of each time node is obtained; and based on the timestamp of each frame of the infrared image, the corresponding position data is matched with the image to form a "high-precision position" associated position sequence.

[0062] according to Figure 6 The process described first receives standardized data from the upstream spatiotemporal synchronization module, including a sequence of infrared images with timestamp synchronization completed, and precise geographic location information for each frame reconstructed through TMS anchor point calibration and Doppler velocity integration. Then, based on preset parameters, the target sequence specifications are calculated, given the effective shooting distance D = 8 meters and the target's normalized vehicle speed V. target =80km / h (≈22.2m / s), infrared frame rate FPS=70Hz, first calculate the time it takes for the component to pass through an 8-meter field of view. Thus, the length of the target sequence is obtained:

[0063] Each frame, i.e., the image sequence within an 8-meter field of view at all vehicle speeds, needs to be standardized to 25 frames; subsequently, different processing strategies are adopted based on the difference between the real-time vehicle speed and the target normalized vehicle speed:

[0064] Low-speed scenario (e.g., 40km / h ≈ 11.1m / s): Original travel time T low =8 / 11.1≈0.72s, original frame count N low =0.72×70≈50 frames (original frame count > target frame count). Using the equidistant frame extraction method, 25 frames are uniformly extracted at intervals of 0.32m / frame within an 8-meter field of view, and redundant data is removed.

[0065] Average speed scene (80km / h≈22.2m / s): The original frame count is exactly 25 frames, so the original sequence is used directly;

[0066] High-speed scenario (e.g., 120km / h ≈ 33.3m / s): Original travel time T high =8 / 33.3≈0.24s, original frame count N high =0.24×70≈17 frames (original number of frames < target number of frames). Using the equidistant frame interpolation method, new frames are inserted based on the motion compensation technology of adjacent frames to complete the sequence to 25 frames.

[0067] The comparison of the number of frames in the image sequence before and after normalization at different vehicle speeds is shown in Table 1 (before normalization) and Table 2 (after normalization):

[0068] Table 1: Original Frame Count Dependence on Vehicle Speed ​​Before Normalization

[0069]

[0070] Table 2: Standard sequence list after normalization to 25 frames

[0071]

[0072] Finally, the normalized sequence is quality verified. Spatial consistency is evaluated by calculating the standard deviation of the spatial sampling interval (the consistency score is required to be no less than 0.9). Motion compensation technology is used to maintain temporal continuity and avoid image blurring caused by interpolation. This ensures that the spatial interval of the output sequence is stable at 0.32±0.02m / frame, the processing delay is controlled within 5 milliseconds, and the information loss rate is less than 3%.

[0073] Step 3: Object detection and cross-frame tracking:

[0074] Contact line components are identified from the normalized image sequence, and continuous temporal trajectories are formed through cross-frame correlation, providing a unified analysis object for subsequent spatiotemporal feature fusion, such as... Figure 4 As shown in the figure, this image demonstrates the detection effect of contact wire components in an infrared image. The target detection model clearly locates the bounding boxes of components such as disconnect switches, upper and lower anchor insulators, and dropper clamps, clearly defining the specific location of each component in the image. The tracking process must follow standardized logic; please refer to [reference needed]. Figure 7 The detailed implementation process is as follows:

[0075] For each frame of normalized infrared image, a target detection model based on the YOLOv8 architecture is used for inference analysis. This model has been specifically trained on a large number of infrared images of overhead contact line components and can accurately identify key components such as insulators, clamps, and locators. It outputs the bounding box coordinates, category label, and confidence score for each detected target. A confidence threshold of 0.6 is set to filter out unreliable detection results with confidence scores below the threshold, ensuring that the detection rate of key components meets the preset requirements.

[0076] See Figure 7 First, a motion model is constructed to predict the position of the component. A Kalman filter model is used to describe the target's motion state, and its state vector is 8-dimensional, expressed as: ;

[0077] Where (u,v) are the center coordinates of the component bounding box. The aspect ratio is h, where h is the height. , , , and represent the rates of change of the above parameters, respectively. The prediction equation is:

[0078] ;

[0079] Calculate the predicted state of the current frame component, where, For the k-frame prediction state based on k-1 frame information, For the optimal state estimation of frame k-1, Let be the state transition matrix for k frames; simultaneously calculate the covariance matrix of the predicted states:

[0080] ;

[0081] in, The covariance matrix of the predicted state, The covariance matrix of the optimal estimate for frame k-1. for The transpose of the matrix, Let be the process noise covariance matrix for k frames.

[0082] Data association is performed, and the historical trajectory is matched with the detection result of the current frame by calculating a comprehensive cost function. The comprehensive cost function is as follows:

[0083] ;

[0084] in, , , These are weighting coefficients, with default values ​​of 0.4, 0.4, and 0.2 (the weights can be adjusted to reduce the impact on appearance in scenarios with interference such as sunlight reflection). The motion cost is calculated by determining the Mahalanobis distance between the detection box and the prediction box; To compensate for the appearance cost, 512-dimensional depth features of the detection boxes are extracted through the ReID network, and the cosine distance of the feature vectors is calculated. The overlap cost is calculated by determining the intersection-union ratio (IUU) of the predicted and detected bounding boxes. Finally, the Hungarian algorithm is used to solve for the optimal match on the comprehensive cost matrix, resulting in matched pairs, unmatched detections, and unmatched trajectories.

[0085] The tracking process, managed and state-enhanced, is controlled by a state machine: For successfully matched trajectories, their Kalman filter state is updated with the corresponding bounding box information, and their "unupdated frame count" (age=0) is reset, while their "hits" are increased; for unmatched bounding boxes, if their confidence is higher than 0.7, they are initialized as new transient trajectories (state is Tentative), with age=0 and hits=1; for existing unmatched trajectories, their age is increased, and if age exceeds the maximum threshold (e.g., 30 frames), the target is considered to have left the field of view or is continuously lost, and the trajectory is terminated; when the consecutive hits of a transient trajectory are ≥3, its state changes to Confirmed, and only confirmed trajectories are output to subsequent modules. Ultimately, a complete spatiotemporal trajectory with a unique ID, containing continuous position and image information, is generated for each successfully tracked component, providing an accurate and continuous data foundation for subsequent spatiotemporal feature fusion and thermal defect diagnosis.

[0086] Step 4: Spatiotemporal Feature Fusion and Consistency Analysis:

[0087] For each confirmed time-series trajectory, features are extracted and fused from spatial, temporal, and environmental dimensions to construct a comprehensive feature set reflecting the thermal state of the component. This step needs to be carried out in conjunction with the decision-making process; please refer to [reference needed]. Figure 8 The detailed implementation process is as follows:

[0088] In the spatial dimension, spatial features are extracted for the target region of the component in each frame of the trajectory: temperature statistical features (average temperature, temperature standard deviation, maximum temperature), temperature gradient distribution (calculated by Sobel operator or gradient operator), target-background contrast features (temperature difference and normalized ratio between the target region and the neighboring region), and high temperature region distribution features (bright spot rate, i.e. the proportion of pixels exceeding the set threshold to the total area of ​​the target region). These features together describe the spatial manifestation pattern of component thermal anomalies.

[0089] In the time dimension, for the component features of continuous frames in the trajectory, time features are extracted, including temperature change trend features (temperature change curve, slope, curvature), temperature stability features (fluctuation index of temperature sequence, autocorrelation), bright spot rate time series features (rising / falling trend and stationarity of bright spot rate change curve), and temperature abrupt change features (detecting abrupt change points in temperature change curve by sliding window method). These features are used to capture the evolution relationship of temperature over time and construct time consistency features.

[0090] By combining the position sequence in the trajectory with a preset scene database, environmental context features are extracted: the environmental type of the component (such as tunnel, station, open area, etc.) is identified, and the corresponding background radiation reference value and ambient temperature are obtained. These features are used to correct the interference of environmental factors on thermal state judgment and provide a basis for subsequent temperature compensation and adaptive threshold adjustment.

[0091] See Figure 8 The method integrates spatial, temporal, and environmental context features: First, all features are normalized (using Min-Max normalization or Z-Score standardization) to eliminate differences in the dimensions and numerical ranges of different features. Then, a feature selection method based on analysis of variance (ANOVA) is used to calculate the contribution of each feature to the judgment of thermal defects, eliminating redundant features with contributions below a threshold. Next, the normalized spatial, temporal, and environmental features are combined in a preset order using feature concatenation to form a high-dimensional fused feature vector, which is simple to operate and retains all feature information. Finally, the consistency of the fused feature vector is checked. If the correlation between a feature component and other components is too low, the feature extraction process is re-examined to ensure the rationality of the fused features.

[0092] Step 5: Thermal Defect Identification and Decision Output:

[0093] Based on the fusion features and spatiotemporal consistency constraints, thermal defects are identified, and the final detection result is output. This step requires relying on... Figure 8 The decision-making logic shown is implemented as follows:

[0094] The fused spatiotemporal feature vector is input into a pre-trained Lightweight Gradient Boosting Tree (LightGBM) classification model for inference calculation. This model outputs a thermal defect probability value between 0 and 1, representing the likelihood of a real thermal defect in the component. The decision model can also select an anomaly detection model (such as Isolation Forest, One-Class SVM, or autoencoder) based on the sample situation. This is suitable for scenarios lacking thermal defect samples. By learning the feature distribution of normal components, features deviating from this distribution are judged as anomalies. In practical applications, an ensemble decision-making approach using multi-model fusion can also be adopted, combining the outputs of multiple different types of decision models (voting method or weighted average method). If the judgment results of different models conflict, a backup decision logic is invoked to ensure reliability.

[0095] The output of the decision model is validated a second time by combining spatiotemporal consistency constraints: spatial consistency validation needs to check whether high-temperature regions stably appear at specific physical locations of components (such as conductive connection parts) and whether the spatial distribution is concentrated, while verifying the stability of target-background contrast (fluctuation less than a preset threshold); temporal consistency validation needs to analyze whether the temperature curve shows a stable upward trend or a continuous high temperature trend, and whether the degree of fluctuation is low (temperature change standard deviation is less than a preset threshold) to ensure that thermal defects have temporal continuity.

[0096] Based on the probability values ​​and the results of dual consistency verification, the system performs a three-level decision: if the thermal defect probability is higher than 0.8 and both the spatial and temporal consistency scores are higher than 0.7, the component is determined to be a "real thermal defect," the system triggers a level one alarm, and records detailed information (component category, location, thermal defect parameters, etc.); if the thermal defect probability is lower than 0.3, or the temporal consistency performance is extremely poor (such as drastic temperature fluctuations), it is determined to be a "transient interference" (such as solar reflection), and is eliminated without generating an alarm; for cases in the fuzzy range, that is, the probability value or consistency score is at an intermediate level, the system marks them as "requiring further observation" and archives their data for subsequent manual review or in-depth analysis in conjunction with historical data.

[0097] This embodiment achieves intelligent detection of thermal defects in overhead contact system components through precise synchronization of multi-source data, vehicle speed normalization processing, target detection and cross-frame tracking, spatiotemporal feature fusion, and spatiotemporal consistency decision-making.

[0098] Significantly reduces the false alarm rate caused by environmental interference. Through time-series consistency analysis and spatiotemporal constraint verification, it successfully identifies and eliminates instantaneous thermal interference such as solar reflection, avoiding misjudging non-thermal defects as faults. Improves detection stability under different vehicle speed conditions. Through vehicle speed normalization processing, it maintains stable detection performance over a wide speed range, completely solving the problem of detection performance fluctuation caused by train speed changes. Optimizes the utilization efficiency of multi-source data. Through a hybrid synchronization scheme of "high-frequency speed integration + low-frequency position anchor point", it solves the problem of mismatch between TMS update frequency and infrared video frame rate, improving the accuracy of multi-source data fusion. Enhances the robustness and accuracy of thermal defect detection. By integrating spatiotemporal multidimensional features and environmental information, combined with continuous trajectory analysis formed by target tracking, it can ensure the effective detection of real thermal defects while controlling missed detections, making it suitable for complex and diverse detection environments.

[0099] The system described in this embodiment has been deployed and verified on-site in multiple railway bureaus, and has successfully provided early warnings for a large number of real thermal defect hazards, effectively avoiding potential train operation safety accidents. It has significant practical application value and broad prospects for promotion in improving railway power supply safety and realizing intelligent operation and maintenance.

[0100] Example 3

[0101] A method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization is provided, specifically including the following steps:

[0102] S1 Data Acquisition and Multi-Source Information Acquisition:

[0103] This step acquires multi-source data required for the inspection of overhead contact line components, providing fundamental data support for subsequent spatiotemporal synchronization, vehicle speed normalization, and thermal defect identification. Specifically, it includes the following sub-steps:

[0104] S1.1 Overhead Contact Line Component Image Sequence Acquisition:

[0105] Image sequences of overhead contact line components are continuously acquired using vehicle-mounted acquisition equipment. This equipment primarily employs infrared imaging, which accurately captures temperature distribution information on the component surface. Since temperature anomalies are a key indicator of thermal defects, the infrared image sequence directly reflects changes in the thermal state of the overhead contact line components. During acquisition, the vehicle-mounted acquisition equipment must maintain a preset angle with the train's direction of travel to ensure the lens stably covers critical component areas of the overhead contact line, avoiding image omissions due to train vibrations or turns.

[0106] S1.2 High-Frequency Speed ​​Data Acquisition:

[0107] The train's high-frequency speed data is continuously acquired by an onboard speed detection device, which preferably employs a Doppler odometer. The Doppler odometer can sense real-time changes in the rotational speed of the train's wheel axles and calculate high-frequency speed data by combining this data with preset wheel diameter parameters. Its data update frequency is much higher than that of the train's position-speed system, accurately reflecting instantaneous changes in train speed and providing high-frequency data support for subsequent speed normalization processing, including speed interpolation and position sequence generation.

[0108] S1.3 Low-frequency position and velocity information acquisition:

[0109] The train's low-frequency position and speed information are obtained through a train position and speed device, which preferably uses a train management system (TMS). The TMS collects the operating data of each carriage through the train bus and periodically broadcasts the train's absolute position (such as the mileage markers along the line) and average speed information. Its position information has absolute reference value and can be used as an "anchor point" for the subsequent position sequence generation, correcting the position accumulation error that may occur during the integration of high-frequency speed data.

[0110] S2 Spatiotemporal Synchronization and Vehicle Speed ​​Normalization:

[0111] This step eliminates the impact of time deviations and vehicle speed variations on detection by unifying the time of multi-source data, generating location sequences, and normalizing vehicle speed. This provides standardized data for subsequent target detection and time-series analysis. Specifically, it includes the following sub-steps:

[0112] S2.1 Multi-source data time synchronization:

[0113] Based on a time synchronization protocol, the timestamps of image sequences, high-frequency speed data, low-frequency location information, and low-frequency speed information are processed uniformly. First, an independent time acquisition unit is allocated to each data source to record the original timestamp (accurate to the microsecond level) when each data point is generated. Then, using the time of the train TMS system as the reference clock (the TMS system time is synchronized with the train dispatch center and has global consistency), the timestamps of other data sources are adjusted through the time synchronization protocol to align all data in the same time coordinate system.

[0114] For example, if the timestamp of the infrared image acquisition device lags behind the TMS reference time by Δt, then the timestamps of all image sequences acquired by the device are uniformly increased by Δt to ensure that each frame of image, each velocity data point, and each location information corresponds to a unique and consistent time node.

[0115] S2.2 High-precision position sequence generation:

[0116] By combining low-frequency location information with high-frequency velocity data, a high-precision location sequence corresponding to the image sequence is generated. The specific process is as follows: First, determine the location anchor points: extract the low-frequency location information broadcast by the TMS system, record the timestamp corresponding to each location information, and form a "time-location" anchor point pair;

[0117] The second step is velocity interpolation supplementation: within the time interval corresponding to two adjacent anchor points, high-frequency velocity data is used for interpolation to obtain the continuous velocity value of each time node in the interval. The interpolation method can be selected according to the velocity fluctuation situation, either linear interpolation (suitable for scenarios with stable velocity) or cubic spline interpolation (suitable for scenarios with large velocity fluctuation).

[0118] The third step is to calculate the position by integrating the speed: the interpolated continuous speed signal is integrated over time intervals to calculate the position offset of each time node relative to the previous anchor point. Combined with the absolute position of the previous anchor point, the absolute position of each time node is obtained. The fourth step is to align the position with the image: based on the timestamp of each frame of the image sequence, the corresponding position data is matched from the calculated "time-position" sequence to form a high-precision position sequence that corresponds one-to-one with the image sequence, ensuring that each frame of the image can be associated with the precise position of the train at the time of acquisition.

[0119] S2.3 Image sequence vehicle speed normalization processing:

[0120] Based on vehicle speed information, image sequences at different vehicle speeds are mapped to an equidistant sampling spatial scale, achieving normalization of frames acquired within the same physical distance. The specific process is as follows:

[0121] The first step is to set normalization parameters: preset target vehicle speed (determined based on the common operating speed range of trains and the frame rate of infrared image acquisition equipment) and target number of frames (calculated based on the preset effective detection distance and target vehicle speed to ensure that the number of frames acquired for the same component within the effective detection distance is consistent).

[0122] The second step is to determine the relationship between the current vehicle speed and the target vehicle speed: If the current vehicle speed is lower than the target vehicle speed, the original number of frames in the image sequence will be more than the target number of frames (because the vehicle speed is slower, the component stays within the effective detection distance for a longer time, and more frames are collected). In this case, an equidistant frame extraction method is used, which determines the frame extraction interval based on the effective detection distance and the target number of frames, and uniformly extracts the target number of frames from the original image sequence, removing redundant data; If the current vehicle speed is higher than the target vehicle speed, the original number of frames in the image sequence will be less than the target number of frames (because the vehicle speed is faster, the component stays for a shorter time, and fewer frames are collected). In this case, an equidistant frame interpolation method is used, which generates missing image frames based on the image information (such as component position and temperature distribution) and position sequence of adjacent frames, and fills the sequence to the target number of frames; If the current vehicle speed is equal to the target vehicle speed, the number of frames in the original image sequence is exactly the same as the target number of frames, and the original sequence is directly used for subsequent processing.

[0123] S3 target detection and cross-frame tracking:

[0124] This step identifies contact wire components in the image through target detection, and combines motion prediction and cross-frame correlation to form the continuous temporal trajectory of the components, providing a unified analysis object for subsequent spatiotemporal feature extraction. Specifically, it includes the following sub-steps:

[0125] S3.1 Target Inspection of Overhead Contact Line Components:

[0126] A deep learning-based target detection model is used to detect catenary components in each frame of a normalized image sequence. First, the normalized infrared image is input into a pre-trained target detection model. The model extracts deep features from the image using convolutional and pooling layers, and then processes the feature map using a detection head to output the bounding box coordinates (determining the component's position in the image), category label (e.g., disconnector, insulator, dropper clamp), and confidence score (reflecting the reliability of the detection result) for each catenary component. Then, a confidence threshold is set to filter out detection results with confidence scores below the threshold, retaining high-reliability component detection information—these components are key areas in the catenary system prone to thermal defects, and accurate location is a prerequisite for subsequent thermal defect identification.

[0127] S3.2 Component motion prediction information calculation:

[0128] Motion prediction information for overhead contact line components is calculated using a motion model, with a Kalman filter model being the preferred approach. The Kalman filter model describes the motion state of the components using an 8-dimensional state vector. This vector includes the center coordinates (u, v) of the component's bounding box, its aspect ratio (γ), height (h), and their respective rates of change. —The center coordinates reflect the position of the component in the image, the aspect ratio and height reflect the shape of the component, and the rate of change reflects the speed of the component's movement and shape change.

[0129] The model is divided into two stages: prediction and update. In the prediction stage, the state of the component in the current frame (including position, shape and rate of change) is predicted based on the optimal state estimate and state transition matrix of the component in the previous frame. In the update stage, after the target detection in the current frame is completed, the predicted state is corrected based on the detection results to obtain the optimal state estimate of the component in the current frame, which provides a basis for motion prediction for subsequent cross-frame association.

[0130] S3.3 Cross-frame target association and temporal trajectory formation:

[0131] Based on the spatial location and motion prediction information from the detection results, a tracking algorithm is used to associate targets across frames, forming a set of temporal trajectories for the same overhead contact line component. The specific process is as follows:

[0132] The first step is to construct the association cost matrix: calculate the association cost between each historical trajectory (the trajectory of the confirmed component in the previous frame) and each detection result in the current frame. The association is based on three types of indicators: motion cost (the consistency of motion state between the detection box and the prediction box is calculated by Mahalanobis distance; the smaller the Mahalanobis distance, the higher the motion consistency), appearance cost (512-dimensional depth features of the components in the detection box are extracted by the ReID network, and the cosine distance of the feature vectors is calculated; the smaller the distance, the higher the appearance similarity), and overlap cost (the degree of spatial overlap between the prediction box and the detection box is calculated by intersection-over-union ratio (IoU); the larger the IoU, the higher the spatial consistency). The three types of indicators are weighted and summed according to preset weights to obtain the association cost matrix.

[0133] The second step is to find the optimal match: the Hungarian algorithm is used to solve the association cost matrix to find the optimal matching pair between the historical trajectory and the detection result of the current frame, so as to realize cross-frame association;

[0134] The third step is trajectory management: For successfully matched detection results, they are assigned to the corresponding historical trajectory, and the trajectory status information (such as current position, motion parameters, and number of consecutive matches) is updated; for unmatched high-confidence detection results (confidence higher than a preset threshold), a new transient trajectory is initialized; for unmatched historical trajectories, if the number of consecutive unmatched frames does not exceed a preset threshold, the trajectory is retained and the prediction for the next frame is performed; if the threshold is exceeded, it is determined that the component has left the detection field of view, and the trajectory is terminated; The fourth step is trajectory status upgrade: The newly initialized trajectory is in a "transient" state, and it needs to be matched successfully for multiple consecutive frames (such as 3 consecutive frames) before it is converted to a "confirmed state". Only the confirmed state trajectory will be output to the subsequent spatiotemporal feature extraction steps to avoid false trajectories caused by temporary interference (such as instantaneous reflection) from entering the subsequent process.

[0135] S4 Spatiotemporal Feature Fusion:

[0136] This step extracts spatial, temporal, and environmental context features for each confirmed time-series trajectory. Feature fusion is then used to form a comprehensive spatiotemporal fusion feature reflecting the thermal state of the component, providing feature support for subsequent thermal defect decisions. Specifically, this includes the following sub-steps:

[0137] S4.1 Spatial Feature Extraction:

[0138] In the spatial dimension, for the target region (determined by the detection box) of each frame image in the trajectory, features reflecting the spatial distribution of the component's thermal state are extracted, specifically including:

[0139] First, temperature statistical characteristics are calculated, including the average temperature of all pixels in the target area (reflecting the overall temperature level of the component), temperature variance (reflecting the uniformity of temperature distribution in the area; a large variance indicates significant local temperature differences, which may indicate thermal defects), and maximum temperature (locating the point with the highest temperature, which is usually the core area of ​​the thermal defect).

[0140] Second, temperature gradient characteristics: the spatial gradient distribution of temperature within the target area is calculated using the Sobel operator or gradient operator. Areas with larger gradients correspond to parts with drastic temperature changes and are potential concentration areas of thermal defects.

[0141] Thirdly, target-background contrast features are used to select the neighboring pixel area around the target area (such as an area with a preset width extended from the detection box), and calculate the difference or ratio between the average temperature of the target area and the average temperature of the neighboring area. A high contrast indicates that the component temperature is significantly higher than the surrounding environment, and there is a greater possibility of thermal defects.

[0142] Fourth, the distribution characteristics of high-temperature areas are analyzed. A temperature threshold is set (determined based on the normal operating temperature range of the component), and the proportion of high-temperature pixels above the threshold to the total area of ​​the target area is statistically analyzed (i.e., the bright spot rate). A high bright spot rate indicates that the influence range of the thermal anomaly area is large.

[0143] S4.2 Temporal Feature Extraction:

[0144] In the time dimension, features reflecting the evolution of the component's thermal state over time are extracted from the target region features of consecutive frames in the trajectory, specifically including:

[0145] First, the temperature change trend characteristics are analyzed by arranging the average temperature of the target area in each frame of the trajectory in chronological order to form a temperature change curve. The slope of the curve is calculated by linear fitting (a positive slope indicates a temperature increase, and a larger absolute value of the slope indicates a faster temperature change). The acceleration of temperature change is analyzed by the curvature of the curve (a positive curvature indicates a faster rate of temperature increase, which may indicate an aggravation of thermal defects).

[0146] Second, temperature stability characteristics. Calculate parameters such as the standard deviation and coefficient of variation (the ratio of standard deviation to mean) of the temperature change curve. Low stability parameters indicate that the temperature change is gradual. If the temperature remains at a high level and the stability is good, it is more likely to be a real thermal defect (rather than a transient disturbance).

[0147] Third, the temporal characteristics of the high-temperature region are analyzed by arranging the bright spot rate of each frame in chronological order to form a bright spot rate change trend curve. The upward / downward trend of the curve is analyzed (e.g., a continuous increase in the bright spot rate indicates that the range of thermal defects is expanding) and the stability is analyzed (e.g., a stable bright spot rate at a high value indicates that thermal defects are stable).

[0148] Fourth, temperature change characteristics: by using the sliding window method to detect abrupt changes in the temperature change curve (such as the temperature in a certain frame increasing by more than a preset threshold compared to the previous frame), the time node corresponding to the abrupt change point can be used as a reference for the sudden occurrence of thermal defects.

[0149] S4.3 Environmental Context Feature Extraction:

[0150] By combining location sequences with a pre-set scene database, contextual features reflecting the detection environment are extracted to correct for interference from environmental factors in thermal state determination. Specifically, these features include:

[0151] First, the environmental type characteristics are defined. A pre-set scenario database stores the geographical range and environmental parameters of different environmental types (such as inside a tunnel, station area, open area, and mountain road section). By matching the location information in the location sequence with the database, the environmental type of the component is determined (e.g., if the location falls within the tunnel mileage range, it is determined to be an environment inside a tunnel).

[0152] Second, background radiation characteristics. The intensity of background radiation varies significantly in different environmental types (e.g., there is no solar radiation in the tunnel, so the background radiation is low; in open areas, the solar radiation is strong at noon, so the background radiation is high). The background radiation reference value corresponding to the current environmental type is queried from the scene database and used to compensate for the temperature value of the target area (e.g., the background radiation value is subtracted from the temperature of the target area to obtain the actual heating temperature of the component).

[0153] Thirdly, the ambient temperature characteristics are collected in real time by the vehicle-mounted environmental sensor, or the average ambient temperature of the current environment type is queried from the scene database. The ambient temperature can be used as a reference benchmark to judge whether the component temperature is abnormal (for example, in a high-temperature environment, the normal operating temperature threshold of the component needs to be appropriately increased to avoid misjudgment due to high ambient temperature).

[0154] S4.4 Spatiotemporal fusion characteristics are formed:

[0155] The extracted spatial features, temporal features, and environmental context features are fused to form a unified spatiotemporal fusion feature vector. The specific process is as follows:

[0156] The first step is feature normalization: Since the units and numerical ranges of different types of features are quite different (e.g., the unit of the average temperature is degrees Celsius, and the range of the bright spot rate is 0-1), Min-Max normalization or Z-Score standardization is used to map all features to the same numerical range (e.g., [0,1]) to avoid fusion bias caused by differences in feature magnitude.

[0157] The second step is feature selection: A feature selection method based on analysis of variance (ANOVA) is used to calculate the contribution of each feature to the judgment of thermal defects, and redundant features with a contribution below the threshold (such as features that are highly correlated with other features) are removed to reduce feature dimensions and improve the inference efficiency of subsequent decision-making models.

[0158] The third step is feature fusion: feature concatenation or attention mechanism fusion methods are adopted. Feature concatenation combines normalized spatial features, temporal features, and environmental features in a preset order to form a high-dimensional feature vector. It is simple to operate and can retain all feature information. Attention mechanism fusion learns the importance weights of different features in different scenarios through a neural network model (e.g., increase the weight of environmental features when there is a lot of environmental interference, and increase the weight of temporal features when there is little temperature fluctuation), dynamically adjusts the feature contribution, and improves the targeting of the fused features.

[0159] The fourth step is feature verification: the consistency of the fused feature vector is checked. If the correlation between a certain feature component and other components is too low (such as the correlation coefficient between environmental features and spatial features being lower than the preset threshold), the feature extraction process is re-examined to ensure the rationality of the fused features.

[0160] S5 thermal defect decision output:

[0161] This step inputs spatiotemporal fusion features into the decision model, combines spatiotemporal consistency constraints for comprehensive judgment, and outputs the final thermal defect detection result to provide decision support for train operation and maintenance. Specifically, it includes the following sub-steps:

[0162] S5.1 Decision Model Reasoning:

[0163] The spatiotemporal fusion feature vector is input into a pre-defined decision model. Based on the judgment rules learned during training, the decision model outputs the probability or category label of the component having a thermal defect. The decision model preferentially adopts a classification model or anomaly detection model: classification models (such as gradient boosting tree LightGBM, support vector machine SVM) are suitable for scenarios with sufficient labeled samples (normal samples and thermal defect samples). The model learns the feature differences between the two types of samples and directly outputs the category label of "normal" or "thermal defect" and the corresponding probability value; anomaly detection models (such as isolated forest, one-class SVM, autoencoder) are suitable for scenarios lacking thermal defect samples. The model learns the feature distribution of normal components and judges features that deviate from this distribution as anomalies (i.e., thermal defects), outputting anomaly probability values.

[0164] S5.2 Spatiotemporal Consistency Constraint Judgment:

[0165] The output of the decision model is validated a second time by incorporating spatiotemporal consistency constraints to eliminate misjudgments caused by transient disturbances. Specifically, this includes:

[0166] First, spatial consistency is assessed, with the core verification being the rationality of the distribution of high-temperature areas and the stability of the target-background contrast. The rationality of the high-temperature area distribution is judged by the matching degree between the center position of the high-temperature area and the functional area of ​​the component (e.g., the conductive connection part of the disconnector is a normal heat-generating area; if the high-temperature areas are concentrated here and compactly distributed, the spatial consistency is high; if the high-temperature areas are scattered in non-conductive parts, the spatial consistency is low). The stability of the target-background contrast is judged by the fluctuation range of the contrast in consecutive frames (e.g., if the contrast remains at a high value and the fluctuation is less than a preset threshold, the spatial consistency is high; if the contrast fluctuates drastically, the spatial consistency is low).

[0167] Second, time consistency is judged, which verifies the stability of temperature changes and the persistence of thermal defects. Temperature change stability is judged by the degree of fluctuation of the temperature change curve (if the temperature does not jump sharply and the slope changes gently, the time consistency is high; if the temperature rises or falls sharply, the time consistency is low). The persistence of thermal defects is judged by the number of frames with high temperature (if the number of frames with high temperature exceeds the preset threshold, the time consistency is high; if high temperature occurs in only a single frame or a few frames, the time consistency is low).

[0168] S5.3 Thermal Defect Judgment Result Output:

[0169] Based on the reasoning results of the decision model and the verification results of the spatiotemporal consistency constraints, the final thermal defect judgment result is output, specifically including three types of results:

[0170] First, there is the "real thermal defect". If the probability of thermal defect output by the decision model is higher than the preset high threshold (e.g., 0.8) and the spatiotemporal consistency verification is passed (both spatial and temporal consistency are higher than the corresponding thresholds), then it is determined to be a real thermal defect, and an early warning message is generated. The content includes component category, location information (from location sequence), thermal defect characteristic parameters (e.g., average temperature, bright spot rate, temperature change slope), and severity level (classified according to temperature level, bright spot rate, and duration). The early warning message is sent to the train operation and maintenance management system through the train communication network, and the trajectory data and characteristic information are stored locally for subsequent maintenance reference.

[0171] Second, "transient interference": if the probability of thermal defects output by the decision model is lower than the preset low threshold (e.g., 0.3), or if the spatiotemporal consistency verification fails (e.g., time consistency is significantly lower than the threshold, or temperature fluctuates drastically), it is judged as transient interference (e.g., solar reflection, temporary external heat source). The event is only recorded in the local log (including the time, location, and characteristic parameters of the interference) and no warning information is generated to avoid false alarms.

[0172] Third, "further observation is needed". If the probability of thermal defects is between high and low thresholds (e.g., 0.3-0.8), or the spatiotemporal consistency verification is partially passed (e.g., high spatial consistency but medium temporal consistency), the trajectory data and feature information of the component will be archived and stored. Subsequently, a secondary analysis will be conducted by combining the results of multiple vehicle inspections (e.g., if the same component shows similar characteristics in multiple vehicle inspections) or historical data (e.g., the normal temperature range of the component). If the secondary analysis confirms the existence of thermal defects, it will be upgraded to "real thermal defects" and an early warning will be generated.

[0173] The detection method described in this embodiment effectively eliminates the impact of time deviation and vehicle speed changes on detection through spatiotemporal synchronization of multi-source data and vehicle speed normalization. Time synchronization ensures the temporal consistency of image, speed, and position data, laying the foundation for multi-source data fusion. Vehicle speed normalization unifies image sequences at different vehicle speeds into a standard temporal sequence, ensuring consistent detection conditions for the same component under different operating states. This effect can be fully achieved through time synchronization, position sequence generation, and vehicle speed normalization in step S2, fundamentally solving the problem of inconsistent temporal features caused by vehicle speed changes in traditional detection methods. By combining target detection with cross-frame tracking, discrete single-frame detection results are transformed into continuous component trajectories, avoiding the limitations of single-frame detection and improving the continuity and accuracy of component identification. Through the fusion of spatiotemporal features and environmental context features, the spatiotemporal evolution of the component's thermal state is comprehensively captured. Combined with a decision-making mechanism constrained by spatiotemporal consistency, it effectively distinguishes between real thermal defects and transient interference (such as solar reflection), significantly reducing the false alarm rate. Meanwhile, the flexible decision-making model and result output mechanism take into account both the real-time nature of detection (real-time early warning generation) and reliability (fuzzy case follow-up analysis), providing efficient and accurate technical means for the intelligent operation and maintenance of the overhead contact system, and effectively improving the robustness and practicality of overhead contact thermal defect detection.

[0174] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization, characterized in that, Includes the following steps: S1. Continuously acquire image sequences of the overhead contact line components using on-board acquisition equipment, and simultaneously acquire high-frequency speed data provided by the on-board speed detection device and low-frequency position and speed information provided by the train position and speed device; S2. Based on the time synchronization protocol, the timestamps of the image sequence, high-frequency speed data, low-frequency position information, and low-frequency speed information are unified. The low-frequency position information and high-frequency speed data are combined to generate a position sequence corresponding to the image sequence. Based on the vehicle speed information, the image sequences at different vehicle speeds are mapped to an equidistant sampling spatial scale to achieve normalization of the acquired frames within the same physical distance. The image sequence is an infrared image sequence. S3. A target detection model is used to detect the catenary components in the normalized image sequence. Combining the spatial location and motion prediction information of the detection results, a tracking algorithm is used to achieve cross-frame target association and form a set of temporal trajectories of the same catenary component. S4. For each time-series trajectory, extract the spatial and temporal features of the target area, and fuse the spatial and temporal features to form a spatiotemporal fusion feature; S5. Input the spatiotemporal fusion features into the decision model, combine them with the spatiotemporal consistency constraints to determine whether there are thermal defects in the contact wire components, and output the thermal defect judgment result; The decision model is an anomaly detection model or a classification model; the spatiotemporal consistency constraints include spatial consistency constraints and temporal consistency constraints, wherein the spatial consistency constraints are the concentration of high-temperature area distribution and the stability of target-background contrast, and the temporal consistency constraints are the stability of temperature change trends and the degree of temperature fluctuation.

2. The method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization according to claim 1, characterized in that, In S1, the on-board acquisition device is an infrared image acquisition device, the on-board speed detection device is a Doppler odometer, and the train position and speed device is a train TMS system.

3. The method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization according to claim 1, characterized in that, In step S2, the low-frequency position information and high-frequency velocity data are combined to generate a position sequence corresponding to the image sequence. Specifically, this includes: using the low-frequency position information as a reference anchor point and the high-frequency velocity data as an interpolation basis, the low-frequency position information is used as a position reference. The velocity change between two adjacent low-frequency position information points is interpolated and supplemented by the high-frequency velocity data. Then, the interpolated velocity signal is integrated to obtain high-precision position data corresponding to each frame of the image sequence, thereby forming the position sequence.

4. The method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization according to claim 1, characterized in that, In step S2, the image sequences at different vehicle speeds are mapped to an equidistant sampling spatial scale based on the vehicle speed information. Specifically, this includes: presetting a target vehicle speed and a target number of frames; if the current vehicle speed is lower than the target vehicle speed and the original number of frames in the image sequence is greater than the target number of frames, an equidistant frame extraction method is used to extract images of the target number of frames from the original image sequence; if the current vehicle speed is higher than the target vehicle speed and the original number of frames in the image sequence is less than the target number of frames, an equidistant frame interpolation method is used to insert images into the original image sequence to make up for the target number of frames; if the current vehicle speed is equal to the target vehicle speed, the original image sequence is used directly.

5. The method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization according to claim 1, characterized in that, In step S3, the target detection model is a deep learning-based target detection model, and the contact wire component includes at least one of disconnecting switches, insulators, and dropper clamps.

6. The method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization according to claim 1, characterized in that, In step S3, the motion prediction information is calculated through a motion model, which is a Kalman filter model; the correlation basis of the tracking algorithm includes at least one of the following: the motion distance between the detection result and the predicted position, the similarity of appearance features between the detected target and the tracked target, and the spatial overlap between the detection box and the prediction box.

7. The method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization according to claim 1, characterized in that, In S4, the spatial features include at least one of the following: the average temperature of the target area, the temperature variance, the temperature gradient distribution, the contrast between the target area and the neighboring pixel areas, and the area ratio of the high-temperature area to the target area.

8. The method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization according to claim 1, characterized in that, In S4, the temporal features include at least one of the following: the temperature change curve of the target region in consecutive frames, the temperature change slope, the temperature stability index, and the temporal change trend of the area ratio of high-temperature regions.

9. The method for detecting thermal defects in overhead contact system components based on spatiotemporal fusion and vehicle speed normalization according to claim 1, characterized in that, In step S4, when integrating the spatial and temporal features, environmental context features are also incorporated. The environmental context features are obtained by combining the location sequence with a preset scene database and include at least one of the following: the environmental type of the contact wire component, the ambient background radiation value, and the ambient temperature.

Citation Information

Patent Citations

  • Space-time fusion multi-target tracking method, device, equipment and medium

    CN117314965A

  • Light guide plate defect detection method and system based on neural network

    CN121033057A