A device for conducting simultaneous inspection of overhead line safety patrol based on image recognition
By using coaxially mounted visible light and infrared cameras, combined with a mirror grid sequence and a progressive depth focus strategy, efficient, real-time multimodal detection of the contact network is achieved, solving the problems of insufficient synchronous acquisition and detection accuracy in existing technologies and improving detection efficiency and reliability.
Patent Information
- Application Number
- CN202511066083.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Existing contact network inspection technology has difficulty achieving simultaneous acquisition of visible light and infrared images during high-speed train operation, resulting in the inability to timely detect early thermal hazards such as internal temperature rise in the conductors and abnormal joint resistance. In addition, cameras in different spectral bands have difficulty achieving pixel-level correspondence under high-speed movement, affecting detection accuracy and efficiency.
Coaxially mounted visible light and infrared cameras are used to achieve synchronous image acquisition and a progressive depth focus strategy for the mirrored grid sequence. Combined with multi-feature fusion and cross-segment consistency comparison, a set of suspected defects is generated for manual review.
It has achieved high-precision, real-time detection of the contact network during high-speed train operation, significantly improved the sensitivity of abnormality detection and image clarity, reduced the intensity of manual maintenance, and increased the operation and maintenance safety margin and train power supply reliability.
Smart Images

Figure CN120543959B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of railway inspection, and in particular relates to a device for conducting simultaneous photographing and inspection of overhead line safety inspection based on image recognition. Background Art
[0002] The catenary is a key component in high-speed railway power supply systems, performing both electrical and electrical functions. Its structure is constantly exposed to wind, rain, vibration, arc erosion, and temperature cycling, making it susceptible to defects such as wear, loosening, foreign matter adhesion, and hard spot erosion. To ensure stable sliding contact between the train pantograph and the conductors, the industry has developed a variety of inspection technologies. From the earliest manual walking inspections and binocular spot checks, to the later use of portable SLR cameras and laser rangefinders, and finally to today's ubiquitous catenary inspection vehicles, these technologies have indeed contributed to improving coverage and detection accuracy. However, as train operating speeds have increased from 200 km / h to 350 km / h, traditional technical solutions have gradually exposed significant limitations.
[0003] Existing inspection vehicles are typically equipped with high-speed linear or area scan cameras. They use white-light flashes and laser profilometers to collect conductor geometry and wear information, which is then analyzed offline using back-end software. However, linear scan cameras require the train's speed to match the camera's to maintain pixel aspect ratio, otherwise distortion will occur. While area scan cameras can capture a two-dimensional image, they are susceptible to motion blur due to long exposure times at speeds of 300 kilometers per hour. Furthermore, most inspection vehicles only collect visible light information, making it difficult to detect early thermal hazards such as internal conductor temperature rise and abnormal joint resistance. In recent years, some research institutes have experimented with adding infrared thermal imagers or millimeter-wave radars to their inspection vehicles to achieve multimodal fusion inspection. However, in environments characterized by high speeds, severe vibration, and wide variations in illumination, cameras in different spectral bands often struggle to achieve strictly synchronized pixel-level alignment due to trigger delays and misaligned optical axes. This results in coarse-grained comparisons in the back-end algorithm, which falls short of the precise defect identification required using deep learning. Summary of the Invention
[0004] In view of this, the main purpose of the present invention is to provide a contact network safety inspection and inspection device based on image recognition.
[0005] The technical solution adopted in the present invention is as follows:
[0006] A device for conducting safety inspection of contact network based on image recognition, comprising: an acquisition terminal, an intelligent analysis terminal and an observation tablet; the acquisition terminal is fixedly mounted on a train, and synchronously acquires continuous visible light images and infrared images of the contact network during the operation of the train, and constructs an original multimodal data stream queue in chronological order in a local cache, divides the original multimodal data stream queue into equal-length data segments according to a preset frame group rule, and transmits the data to the intelligent analysis terminal in real time via a high-speed, low-latency link; the intelligent analysis terminal calls a mirrored progressive multi-fusion defect inference algorithm, and performs the following specific process on each data segment: generates a mirrored grid sequence The visible light image and the infrared image are grid-matched in the virtual mirror plane; a progressive depth focus strategy is adopted to realize focal plane sliding grid by grid in the mirror grid sequence, and a clear focus frame is extracted after each sliding; multi-feature fusion is performed on the clear focus frame, and the fusion result is included in the defect candidate set; the inference and judgment process is executed on the defect candidate set, and the inference and judgment process performs cross-segment consistency comparison on the candidate set according to the segment index table, and outputs the suspected defect set and the corresponding segment identifier; the suspected defect set and the corresponding segment identifier are synchronously pushed to the observation tablet, and the inspection personnel manually review the suspected defect set through the observation tablet.
[0007] Furthermore, the acquisition end is provided with a visible light camera and an infrared camera, which are coaxially installed and share the optical axis center, and the optical center distance between the two does not exceed 30 mm; the single-frame spatial resolution of the visible light camera is not less than 1920×1080 pixels, and the frame rate is not less than 60 frames per second; the detection band of the infrared camera is 8 microns to 14 microns, the single-frame spatial resolution is not less than 640×512 pixels, and the frame rate is not less than 60 frames per second.
[0008] Furthermore, the acquisition end writes a unified time stamp for each frame during the acquisition process of visible light images and infrared images, and uses a phase-locked loop and train speed pulse signal for synchronous calibration to keep the inter-frame time difference between the two types of images within 1 millisecond; when the train speed exceeds 200 kilometers per hour, the acquisition end automatically increases the visible light camera frame rate to 120 frames per second and enables spatial four-level downsampling differential compression to ensure that the actual bandwidth occupied by the link does not exceed 80% of the rated bandwidth.
[0009] Furthermore, the process of generating a mirror grid sequence by the intelligent analysis end includes: taking the intersection of the optical axes of the two cameras as the plane origin, arranging seed points in the virtual mirror plane according to the radial spiral method with the golden angle as the increment, the number of seed points is not less than 800, and the plane distance between any two adjacent seed points is maintained between 6 mm and 8 mm; using the Thiessen polygon segmentation method based on the seed points to form a honeycomb mirror grid sequence, and the area variance of a single grid in the mirror grid sequence is controlled within 5% to ensure that the area of each grid is approximately consistent; during the operation of the train, all seed points are rotated 137.5 degrees according to the golden angle and translated 1 pixel in the direction of the train movement every 5 frames, and a new mirror grid sequence is generated in real time, thereby forming a time-progressive mirror grid sequence.
[0010] Furthermore, the mirror grid sequence ensures complete coverage of the virtual mirror plane within any 60 consecutive frames; the intelligent analysis end performs perspective mapping on the visible light image and infrared image acquired at the same time to the current mirror grid sequence respectively. Sub-pixel precision bilinear interpolation is used in the mapping process to uniquely attribute each pixel to the corresponding grid.
[0011] Furthermore, before the train departs, the acquisition end performs an autofocus process on the visible light camera and the infrared camera respectively, determines the optical zero-point focal length of the train when it is stationary, and writes the focal length into the control register as the focal length reference for the entire process; the intelligent analysis end arranges the traversal list grid by grid in ascending order of index, starting from the grid with the smallest index according to the grid index of the current mirrored grid sequence; when traversing to the grid with the largest index, it immediately backtracks in the opposite direction to form a round-trip loop sequence, ensuring that each grid is visited twice in a continuous traversal cycle and the visit interval time is balanced; for each grid in the traversal list, read the image distance increment value recorded when the grid was visited last time, and call the linear interpolation table to calculate the new image distance increment value in real time in combination with the current speed of the train, vibration amplitude and ambient light brightness, and the new image distance increment value is obtained. The distance increment value is used as the step size of this focal plane sliding, and the step size value range is limited to within ±0.20 mm in the optical axis direction; when the traversal process enters a certain target grid, the electric focus mechanism of the visible light camera and the infrared camera is driven to move the focal plane along the optical axis at a uniform constant speed, and the moving distance is equal to the newly added image distance increment value; after the focal plane moves, it remains stationary for not less than 0.50 milliseconds to eliminate the coupling effect of mechanical inertia and vehicle body vibration; in the time slot when the focal plane completes the sliding and remains stationary, the Laplace transform is called to calculate the full-frame sharpness score of the current visible light image and infrared image respectively; the full-frame sharpness scores of the two images are weighted averaged, and the resulting comprehensive sharpness score is compared with the comprehensive sharpness score when the grid was last visited. If the improvement is not less than 5%, the current focus is determined to be valid.
[0012] Furthermore, when the intelligent analysis end determines that the current focus is valid, it marks the current visible light image and the infrared image as a clear focus frame and writes them into the cache; at the same time, it records the current grid index, focus timestamp, comprehensive sharpness score and image distance increment value, and updates the access record of the corresponding grid; for any two clear focus frames with the same grid index and an access time interval of less than 60 frames in the same mirrored grid sequence, a phase consistency comparison is performed; if the comparison results are lower than the threshold in the three indicators of grayscale mean square error, infrared radiation mean difference and edge gradient difference, the clear focus frame with a newer timestamp is retained and the older clear focus frame is invalidated to reduce redundant data.
[0013] Furthermore, within each pair of in-focus frames, the intelligent analysis end calculates the following components for the visible light image and infrared image with the same raster index: visible light brightness mean, visible light gradient directional distribution mean, visible light edge density, infrared radiation intensity mean, infrared radiation gradient directional distribution mean, infrared temperature difference amplitude, cross-spectral histogram, cross-spectral phase consistency, and focal plane image distance increment. These nine components are concatenated in sequence to obtain a raster multi-component descriptor of length 9, which is stored in 32-bit floating-point format. For the same mirrored raster index, no fewer than 60 raster multi-component descriptors are collected within the first 120 frames after the train starts. The mean and standard deviation of each component are calculated and normalized. The normalized raster multi-component descriptors are input into a three-layer fully connected network with 27, 18, and 9 nodes per layer, respectively. The activation function uses a rectified linear unit with a leakage coefficient of 0.01. The output vector and the raster index together form a fusion vector, which is written to the candidate buffer queue in real time.
[0014] Furthermore, for each fusion vector, the intelligent analysis terminal queries the reference vector library of the corresponding raster index; the reference vector library is updated online with an exponential decay coefficient of 0.95. The distance between the current fusion vector and the reference vector is measured, and the distance threshold is preset to 2.5 times the historical standard deviation of the reference vector; if the distance exceeds the threshold, the fusion vector is registered as a potential anomaly and written into the defect candidate set together with the raster index, timestamp and distance value. The defect candidate set adopts a double-ended queue structure and appends new entries in chronological order; once the set size exceeds 1200, the oldest entry is removed according to the first-in-first-out principle; the inference judgment process is triggered immediately when each potential anomaly is added; the inference judgment process includes: for entries in the defect candidate set with a time interval of no more than 15 frames and a raster index difference of no more than 3, they are considered to be the same A physical cluster is constructed; the maximum distance value within the same cluster is taken as the cluster distance value; according to the segment index table, the global frame number of each cluster entry is mapped to the segment identifier; if the same cluster appears in two or more consecutive segments, the cluster is marked as cross-segment consistent; the cluster distance value of the cross-segment consistent cluster and the number of segments where it appears are weighted averaged to obtain the confidence level; when the confidence level is higher than 0.8, the cluster entry enters the suspected defect set; for the entries in the suspected defect set with a raster index difference of no more than 2 and the same segment identifier, a union-find algorithm is used to perform a merge operation, and the one with the highest confidence level is retained as the representative entry; each representative entry contains the raster index, segment identifier, confidence level, and first frame timestamp.
[0015] The above technical solution achieves the following beneficial effects: By forming a closed loop with an acquisition terminal, an intelligent analysis terminal, and an observation panel, the present invention implements simultaneous acquisition of visible light and infrared images of the overhead contact network during train operation, mirroring with progressive grid mapping, grid-level deep focusing, multispectral feature fusion, and cross-segment consistency inference. First, coaxial dual-camera, same-frame acquisition avoids parallax misalignment, and the random rotation and translation of the mirrored grid sequence eliminates image smear caused by high-speed displacement, ensuring uniform coverage of any conductor area within a short time window. Second, progressive deep focusing breaks down traditional full-frame focusing into grid-based micro-displacement focusing, and introduces an online interpolation compensation model, making the focal plane adaptive to changes in vehicle speed, vibration, and illumination, ensuring clear frame output even under sustained high-speed conditions. Third, a 9-dimensional multi-component descriptor fuses visible light texture information with infrared thermal field features. This is then mapped into a unified semantic vector using a lightweight neural network. This is then combined with an exponentially decaying background model for outlier detection, reducing reliance on offline annotated data and enabling real-time self-learning with seasonal and environmental drift. Ultimately, cluster merging and cross-segment confidence assessment based on dual-neighbor constraints on the time grid significantly suppress false alarms, with only high-confidence results pushed to the observation tablet for manual review. Compared to existing inspection schemes that rely on single spectra, low frame rates, or fixed sampling windows, this invention achieves synergistic improvements in image clarity, anomaly detection sensitivity, real-time performance, and bandwidth utilization efficiency. It can continuously output a stable and reliable list of suspected defects under conditions of high-speed train operation, rapid changes in illumination, and complex vibrations, significantly reducing the intensity of manual secondary maintenance and improving the safety margin of overhead line maintenance and train power supply reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of the structure of a device for conducting simultaneous inspection and photo-taking of overhead line safety inspections based on image recognition provided by an embodiment of the present invention;
[0017] Figure 2 Schematic diagram of performance comparison experimental results of the mirroring progressive multi-element fusion defect inference algorithm provided by an embodiment of the present invention;
[0018] Figure 3 This is the working principle of the progressive depth focus strategy provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] All features disclosed in this specification, or all steps in the disclosed methods or processes, except mutually exclusive features and / or steps, can be combined in any manner.
[0020] Any feature disclosed in this specification (including any appended claims and abstract), unless otherwise stated, may be replaced by other equivalent or similar features. In other words, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.
[0021] refer to Figure 1 , a device for conducting safety inspection of overhead line based on image recognition, comprising: a collection terminal, an intelligent analysis terminal and an observation tablet;
[0022] The acquisition terminal is fixedly mounted on the train. The overall framework is built around the synchronous acquisition of continuous visible light and infrared images of the contact network. Its core consists of a visible light camera, an infrared camera, a time synchronization unit, a cache management unit, a compression encoding module, and a high-speed, low-latency link interface. The visible light camera and the infrared camera are coaxially mounted and arranged to share the center of the optical axis. The optical center distance between the two is kept to a very small range, thereby maintaining consistent field of view overlap during high-speed train operation. Both cameras are connected to a unified trigger control circuit at the acquisition terminal. After receiving the train speed pulse signal, the time synchronization unit first calls the phase-locked loop to achieve frequency capture, and then writes a unified time stamp for each frame to ensure that the time difference between the visible light image and the infrared image is maintained within milliseconds.
[0023] When the train starts, the cache management unit immediately constructs a chronological queue of raw multimodal data streams in local memory. This queue continuously appends new frames and dynamically segments them according to preset frame grouping rules, allowing equal-length data segments to be directly mapped into the processing windows of the subsequent mirrored progressive multi-fusion defect inference algorithm. To alleviate link pressure, the compression encoding module adaptively compresses visible light images using a spatial four-level downsampling differential compression strategy. When the train speed exceeds a specified threshold, the visible light camera frame rate is automatically increased and the compression strategy is simultaneously activated, ensuring that the actual link bandwidth remains below the reserved upper limit of the rated bandwidth. The cache management unit creates an index table for each data segment and records the starting timestamp, enabling the intelligent analysis end to generate a mirrored grid sequence and implement a progressive depth-of-focus strategy. While writing to the local cache, the acquisition end transmits the data segments in real time over a high-speed, low-latency link. The link interface utilizes a redundant design, triggering dual-channel load balancing when the risk of frame drop occurs to maintain data integrity.
[0024] When re-entering a section with ample bandwidth, the cache management unit clears the confirmed received data segment by segment according to the segment index table to avoid storage overflow. In addition, the control firmware of the acquisition end automatically performs an autofocus process on each of the two cameras before the train departs, obtains the optical zero-point focal length of the train when it is stationary, and writes it into the control register. This baseline value is then used by the intelligent analysis end to adjust the image distance increment reference of each grid in the mirrored grid sequence. The acquisition end monitors the ambient light brightness and vibration amplitude in real time. When the brightness drops suddenly or the vibration increases suddenly, the compression encoding module prioritizes reducing the compression ratio and instructs the visible light camera to switch the gain mode to maintain the effectiveness of the focus frame acquisition; at the same time, the time synchronization unit realigns the time stamps of the two images to ensure that the inter-frame consistency of the original multimodal data stream queue is not destroyed. The entire acquisition end is connected to the train body through a modular metal frame. The protective cover uses multi-layer protective glass and is coated with an anti-fouling coating to ensure that the mirror surface is clean under long-term operation.
[0025] The health status of the acquisition end is uploaded to the observation tablet in real time. Inspection personnel can remotely view the equipment temperature, power current, and link occupancy in the train maintenance area. If any indicator exceeds the limit, the system will automatically suspend data transmission and retain the latest several segments of the original multimodal data stream queue in the local circular buffer, and reissue it after the fault is eliminated. Through this serial hardware layout and software process, the acquisition end not only provides a high-speed, low-latency link with the intelligent analysis end, but also ensures the time continuity, spatial consistency, and image distance benchmark required for the generation of the mirrored grid sequence. It lays a reliable source data foundation for the subsequent progressive deep focus strategy to extract focused frames, multi-feature fusion to generate defect candidate sets, and inference judgment process to output suspected defect sets based on the segment index table, thereby realizing the stable operation and efficient data link support of the contact network safety inspection and side-by-side inspection device based on image recognition in high-speed train driving scenarios.
[0026] After the acquisition end sends equal-length data segments into the high-speed, low-latency link, the intelligent analysis end first establishes a mirror grid sequence, using the intersection of the optical axes of the two cameras as the origin of the virtual mirror plane, and using radial spirally distributed seed points plus Thiessen polygon segmentation to construct a honeycomb grid. This grid system is equivalent to re-splitting the continuous image into a set of digital sampling units with a fixed topology; then, the seed points are subjected to golden angle rotation and slight shift operations every several frames, so that the mirror grid sequence presents a quasi-random, uniform but predictable displacement trajectory on the time axis, thereby mapping the relative displacement caused by the high-speed movement of the train into a regular and controllable coordinate drift, achieving the purpose of "offsetting the target movement with coordinate movement". This step solves the tracking problem caused by the rapid passing of contact network elements in the image.
[0027] After obtaining unified grid coordinates, the intelligent analysis end aligns the visible and infrared images at the frame level and projects them onto the current mirrored grid sequence. Each pixel has a unique grid affiliation and retains a time stamp, enabling precise alignment of multimodal information at the same spatial granularity. Unlike traditional global focus, the progressive depth focus strategy delegates the freedom of focal plane sliding to the grid dimension: the system traverses the grid index back and forth. At each grid, it retrieves an interpolation table based on the previous image distance increment and real-time environmental conditions such as speed, vibration, and brightness. It calculates the new focal plane step size in real time and drives the electric focus mechanism to perform micro-shifts. After the micro-shift, the Laplace sharpness is evaluated within a millisecond-level static window. The visible and infrared sharpness are weighted and compared with historical values. If the improvement is significant, the frame is considered in focus. This evolution allows the focus operation to evolve from a "one-time focus operation for the entire frame" to "multiple attempts per grid to find the local extremum." Focus adjustment is discretized and minimized in sync with vehicle body vibration, optically reducing the risk of blur caused by high-speed driving.
[0028] After acquiring a clear frame, the multi-dimensional feature fusion module extracts nine complementary components for the same raster: visible light brightness, visible light gradient directional distribution, visible light edge density, infrared radiation intensity, infrared radiation gradient directional distribution, infrared temperature difference amplitude, and three cross-spectral correlation statistics. These components are then combined with the current image distance increment to form a multi-component raster descriptor. Compared to single-spectral analysis, this approach integrates texture, structure, and thermal field information into a single vector space, allowing abnormalities that were previously in different physical domains to be projected onto a common metric. The descriptor is normalized and fed into a three-layer fully connected network to generate a fused vector, which is then written to the candidate buffer along with the raster index. The system maintains a reference vector library for each raster index, using exponentially decaying updates to slowly converge the library vectors to the "normal feature center" of the raster over time. When the distance between a new fused vector and a reference vector exceeds a threshold set by the historical standard deviation of the reference vector, it is marked as a potential abnormal write defect candidate. The core idea here is to transform the anomaly detection problem into a distance judgment based on an "online background model - instantaneous distance from the group" approach, avoiding reliance on prior offline training data and allowing the model to adapt to seasonal changes in light and temperature. The defect candidate set then enters the inference and judgment process: within a dual-dimensional sliding window of time and grid dimensions, the system clusters potential anomalies with small intervals and close indices into physical clusters. The global frame number is then mapped to a data segment identifier using a segment index table. If the same cluster persists across multiple segments, it is considered cross-segment consistent, indicating that the anomaly represented by this physical cluster has been independently observed multiple times during the train's journey, making it more reliable than a transient anomaly. For cross-segment consistent clusters, the algorithm calculates a confidence score by weighting the maximum distance and the number of segments in which it appears. High-confidence clusters are then added to the suspected defect set. A parallel lookup merge algorithm merges representative entries with similar grid indices and identical segment identifiers to prevent the same physical defect from being recorded multiple times due to grid drift.
[0029] The final set of suspected defects is pushed to the observation tablet in chronological order, allowing inspectors to view the corresponding thumbnails and confidence levels. The entire process is essentially a multi-stage nested loop: mirroring the grid sequence solves spatial alignment and temporal uniform sampling, progressive depth focus experimentally searches for the local optimal image distance and avoids motion blur, multi-feature fusion compresses optical and thermal information into the same feature tensor, a reference vector library provides an online self-updating background model, distance metrics enable outlier detection, and cross-segment consistency checks enhance confidence in the temporal dimension. This layered iterative design, from the physical coordinate system, optical system, feature expression, to statistical inference, enables the intelligent analysis end to perform real-time, stable, and threshold-adaptive judgment of minor defects in the contact network under real-world scenarios such as high-speed driving, variable lighting, and severe vibration. High-confidence results are then fed back to the observation tablet via the link interface, laying a reliable foundation for data processing and anomaly inference for image recognition-based contact network safety inspection devices.
[0030] The observation tablet, serving as a human-machine interface terminal, is installed in the train maintenance area. It maintains bidirectional synchronization with the intelligent analysis terminal via a high-speed, low-latency link. Its core function is to present suspected defect sets and corresponding segment identifiers in real time on a portable interface and support manual verification. On the hardware side, the observation tablet utilizes a fully laminated, high-brightness display panel and a multi-point capacitive touchscreen. Its housing is a vibration-resistant and dust-proof alloy frame. The battery module supports continuous operation without an external power source throughout the train's journey. The network interface supports both wired Gigabit Ethernet and the train's onboard 5G private network, enabling automatic hot-standby link switching. The software architecture comprises a message reception service, a data caching layer, an interactive rendering engine, and a feedback upload module. The message reception service continuously listens for suspected defect entries pushed by the intelligent analysis terminal via a multi-threaded queue and writes them to a local circular cache in chronological order. The data caching layer is responsible for secondary index mapping between segment identifier indexes and raster indexes, enabling historical defect retrieval even when offline.
[0031] The rendering engine immediately refreshes the interface upon receiving a new entry. The list area is sorted in descending order of confidence, with each row displaying the raster index, confidence percentage, and segment identifier. Inspectors can tap an entry to expand a multimodal thumbnail and pinch-to-zoom to view a full-resolution comparison of visible and infrared images. A long press enters annotation mode, allowing users to draw polygonal labels or text descriptions on the image. All annotations are timestamped and authenticated. The feedback upload module combines the annotation data with the original entry, encrypts it, and transmits it back to the intelligent analysis client for adaptive model updates. A read-only log is also generated locally on the tablet to prevent record loss. A status bar at the bottom of the interface displays link bandwidth, cache utilization, and battery charge in real time. When any metric approaches a threshold, it automatically enters energy-saving mode, reducing the rendering refresh rate and alerting inspectors to address the issue promptly. To accommodate high-noise environments, the observation tablet is equipped with a dual-microphone array and echo cancellation algorithm, supporting voice wake-up and voice annotation of defect entries. Furthermore, a built-in inertial sensor dynamically adjusts touch sensitivity based on train turbulence to prevent accidental touches. The entire mechanism ensures that no suspected defect information is missed. Inspection personnel can verify the results in a graphical manner at any time and quickly feed back manual judgments to the intelligent analysis terminal, thereby closing the automatic inference and manual confirmation loop of the contact network safety inspection device based on image recognition, and improving the reliability and maintainability of the entire system in high-speed operation scenarios.
[0032] Furthermore, the core imaging unit at the acquisition end consists of a visible light camera and an infrared camera, coaxially mounted on an integrated metal bracket. They share a common optical axis via high-precision locating pins and fine-tuning bolts, and the optical center distance is strictly limited to within 30 mm. This mechanical structure not only compresses the baseline length between the cameras, fundamentally reducing parallax, but also ensures that the overlapping fields of view of the two imaging channels are almost completely identical, laying the physical foundation for grid-level pixel correspondence in the subsequent mirrored grid sequence. The visible light camera selects a single-frame spatial resolution of no less than 1920×1080 pixels, supported by a sampling rate of no less than 60 frames per second, ensuring sufficient temporal resolution to capture minute structural details of the catenary conductors even at high train speeds. Its imaging sensor utilizes a back-illuminated stacking process and a low-noise global shutter, effectively suppressing tic-tac-toe artifacts and rolling smear, ensuring high-quality sharpness scores even after rapidly sliding the focal plane using the progressive depth-of-focus strategy.
[0033] The infrared camera covers the detection band of 8 microns to 14 microns, with a single-frame spatial resolution of no less than 640×512 pixels and a frame rate of no less than 60 frames per second. This band is sensitive to the thermal radiation signals of metal conductors and can reveal early-stage heating anomalies that are difficult to identify in the visible light domain. The clocks of the two cameras are locked to the same time synchronization unit, and their hardware trigger lines use homologous pulses to achieve frame-level synchronous exposure. When the train speed pulse signal fluctuates, the synchronization unit fine-tunes the trigger phase through the phase-locked loop to ensure that the time error between frames is maintained at the millisecond level, meeting the timing accuracy requirements of multimodal fusion. In order to further reduce the imaging error caused by thermal drift, thermal couplers and graphite heat sinks are embedded in the bracket, and the train's running airflow is used to passively cool the components. At the same time, a double-layer optical window is set in front of the camera. The outer layer uses anti-fouling coated glass to resist the impact of high-speed dust, and the inner layer is added with a heating film to prevent frost at low temperatures. The entire imaging unit undergoes a joint calibration process before the train departs. The visible light camera and infrared camera each autofocus and write their optical zero-point focal lengths into the control register. Calibration software then extracts the intrinsic and extrinsic parameter matrices for both imaging paths and stores them in the acquisition-end firmware to ensure geometric consistency during mirrored grid sequence mapping. Short-circuit vibration excitation tests show that the coaxial mounting structure maintains an optical axis offset of less than 0.02 degrees even at a wind speed of 200 km / h. This allows for continuous output of precisely aligned, high-frame-rate, and multi-spectral complementary raw images even in long-term, high-speed operation. This provides a reliable front-end data source for image recognition-based overhead line safety inspection systems.
[0034] Furthermore, when performing synchronous acquisition of visible light and infrared images, the acquisition end first issues a unified hardware trigger pulse to both cameras, followed by a unified time stamp for each frame at the head of the frame buffer. This time stamp is directly derived from a high-precision clock signal generated by a phase-locked loop (PLL), which in turn is locked to the train speed pulse signal in real time. By continuously fine-tuning the phase, the PLL maintains strict alignment of the trigger frequency with vehicle speed changes, converting speed fluctuations into fine-tuning of the timing base, ensuring that the inter-frame time difference between the two image types is always compressed to less than 1 millisecond. To prevent loss of detail due to excessive image displacement at high speeds, the control firmware continuously monitors the train speed register. Upon detecting a speed exceeding 200 kilometers per hour, the visible light camera frame rate is automatically increased from 60 to 120 frames per second, and a four-level spatial downsampling differential compression module is simultaneously activated. This module first performs a four-level pyramid downsampling of the original resolution image, then applies row-column predictive coding to the difference images between adjacent sampling levels. Finally, adaptive threshold zero-block culling is applied to significantly reduce data redundancy while preserving visual detail. The compressed code stream and the infrared code stream are output together through a high-speed, low-latency link. The link management unit dynamically calculates the actual occupied bandwidth. If it detects that the bandwidth is approaching 80% of the rated value, it will feedback to the compression module to adjust the threshold. Otherwise, the original sampling density will be restored step by step. The entire closed loop ensures that the actual occupied bandwidth of the link does not exceed 80% of the rated bandwidth when the frame rate is doubled, while maintaining the accuracy of the unified time stamp and phase-locked loop calibration mechanism, so that the acquisition end can still output time-synchronized, spatially consistent and bandwidth-controllable multimodal data streams under extreme speed conditions, providing reliable input for subsequent mirroring, progressive multi-fusion defect inference.
[0035] Furthermore, the core principle of the mirror grid sequence lies in using a continuously self-updating, yet uniformly densely packed sampling reference system to transform the uncontrollable field of view drift in high-speed motion scenes into controllable coordinate drift. First, a virtual mirror plane is established with the intersection of the two camera optical axes as the plane's origin. Then, a radial spiral method is used to lay out no fewer than 800 seed points on this plane. The radial spiral method uses the golden angle as the increment. Its mathematical properties ensure that new seed points always fall within the sparsest sector adjacent to the previous seed point, thus naturally forming a nearly uniform point distribution without grid constraints. The distance between adjacent points is constrained to between 6 and 8 mm, ensuring that the entire distribution is approximately equidistant at the microscopic scale. Next, the system performs Thiessen polygon segmentation on these seed points. Thiessen polygons divide the plane into a set of non-overlapping, fully overlapping honeycomb-shaped cells, each containing exactly one seed point. Due to the uniform point distribution in the previous step, the segmentation result automatically exhibits approximately uniform area. Furthermore, by limiting the area variance to within 5%, a statistically uniform grid is obtained. The resulting honeycomb grid can be considered a discretized "projection screen" for the camera's field of view. All subsequent pixel mapping, feature accumulation, and focus evaluation are performed on this screen. To accommodate parallax variations caused by the high-speed train movement, the system incorporates a "rotation and translation" mechanism in the temporal dimension: every five frames, all seed points are rotated 137.5 degrees about the origin and translated one pixel along the train's direction of motion. This golden angle rotation ensures that the grid's overall angular position never overlaps with the previous state period, thus eliminating spatial aliasing caused by fixed sampling. The subtle translation allows any physical object to be captured by different grids from slightly offset sub-pixel angles within a short period of time, achieving sub-pixel jitter compensation. Continuously performing this update results in a time-progressive mirrored grid sequence: at the macroscale, it covers the entire field of view and adaptively drifts with motion, while at the microscale, the grid size and distribution remain statistically constant. This sequence provides a balanced target for progressive depth focusing and a unified coordinate benchmark for multivariate feature statistics, enabling subsequent online background learning and outlier detection to be performed in a sampling space that is continuous and uniform in time and space. Ultimately, it ensures that the image recognition-based contact network safety inspection and shooting device can still stably obtain the structural and thermal information required for judgment in high-speed and strong vibration scenarios.
[0036] Furthermore, within the dynamic mechanism of the continuous rotation and translation of the mirrored grid sequence, the system presupposes a key constraint: the total grid coverage area within any consecutive 60 frames must completely cover the virtual mirrored plane, leaving no "sampling shadows." The mathematical significance of this constraint is to couple the field of view division with the time axis. This ensures that even if the high-speed movement of the train causes the target to quickly slide out of a single grid in the image, it will inevitably fall back into the adjacent grid within a very short frame interval, thus ensuring observation continuity. To implement this constraint, the grid generation module not only rotates and translates the seed point during each five-frame update, but also evaluates the cumulative coverage in real time. If the number of grid visits to a particular planar area is detected to be below a threshold, the subsequent rotation step size and translation vector are fine-tuned to ensure that the coverage gap is filled with new grids within the next five frames. At the same time, the intelligent analysis end performs perspective mapping on the visible light image and infrared image acquired at the same time. The core process is to convert the pixel coordinates of the two images into virtual mirror plane coordinates according to their respective internal calibration matrices and external pose matrices. Then, using the current mirror grid sequence as a reference, sub-pixel precision bilinear interpolation is used to calculate the position of each source pixel on the target plane, and uniquely attribute it to a specific grid. The sub-pixel calculation of bilinear interpolation not only compresses the mapping error, but also smoothes the grayscale and radiance values of different source pixels at the grid boundary, so that the grid-level feature statistics are not affected by hard edge roughening; the unique attribution rule eliminates pixel overlap across grids and guides subsequent multi-dimensional feature fusion to be performed at a strict spatial granularity. Ultimately, the full coverage constraint and sub-pixel mapping jointly construct a temporally and spatially consistent multimodal sampling volume, ensuring that each frame provides complete and seamless input for progressive depth focus, online background learning, and outlier detection.
[0037] Furthermore, while the train is stationary and has not yet started, the acquisition end first uses a multi-step micro-shift search to obtain the respective sharp extreme points of the visible light camera and the infrared camera. The fundamental purpose of this process is to establish the absolute coordinate origin for all subsequent dynamic focusing. The system writes the optical zero-point focal length obtained from the search into the control register, so that any subsequent focus adjustments are based on this value, thus forming a stable relative measurement system. After the train starts, the intelligent analysis end no longer judges clarity based on the entire image, but instead treats each grid in the mirrored grid sequence as an independent local observation window. Through a round-trip traversal mode of index increment and reverse backtracking, each grid is accessed twice on the time axis with balanced intervals. This alternating access follows the principle of "local balanced sampling": the goal is to ensure that all spatial units obtain focus opportunities at a similar frequency within limited hardware resources, avoiding convergence deviations caused by continuous omissions. Each time the grid is entered, the system calculates the new image distance increment through an interpolation table based on the image distance increment value left over from the previous visit and the three environmental quantities of speed, vibration, and brightness acquired in real time. The interpolation table is essentially an experience-data mixed mapping function that maps the environmental state to the focus compensation step size. From the perspective of control theory, it is equivalent to an adaptive feedforward link, which estimates the impact of external disturbances on the focal length and offsets them in advance at the execution layer.
[0038] To prevent the focal plane from shifting too far at once and causing loss of control, the step size is limited to ±0.20 mm. This is equivalent to setting a saturation link in the axial channel, ensuring that the electric mechanism always operates in the linear region. After the focal plane moves, a static window of at least 0.50 milliseconds is left. Its fundamental purpose is not to simply wait, but to allow the residual inertia and body vibration generated by the movement to naturally decay with the help of damping. Only when the external force has been basically attenuated can the sharpness be measured to avoid incorporating the instantaneous high frequency of mechanical jitter into the sharpness estimate, ensuring that the evaluation value truly reflects the optical focus quality. The Laplace transform is selected as the evaluation method to calculate the full-frame sharpness because the Laplace operator is most sensitive to the second-order gradient of edges and textures and is invariant to linear scaling of illumination. In multimodal scenarios, the visible light and infrared sharpness scores are weighted averaged. In essence, the two spectral bands complement each other in their response to material details and thermal radiation, reducing the impact of single light source flicker or temperature changes on a single sharpness indicator. The system considers the focus to be valid only when the overall sharpness improves by more than 5% relative to the historical value; the 5% threshold is derived from a large number of experimental statistics. It is located in the dividing range between the instrument's repeated measurement error and the actual focal length improvement, which can both filter out noise and capture the real improvement. If the focus is judged to be invalid, the system does not immediately abandon the grid. Instead, it stores the current environmental quantity and the image distance increment in a short-term buffer, which is used to update the local weights when the interpolation table is resampled, forming a slow and continuous self-learning closed loop. The core idea of this is to use online fine-tuning to gradually approach the optimal compensation curve, rather than a static table lookup, so as to maintain the robustness of the focal length adjustment during long-distance operations with constantly changing vehicle speed, track status, and lighting conditions. During the entire process, round trip traversal provides time-balanced sampling, the interpolation table is responsible for feedforward compensation, the step size saturation and static window ensure stable execution, and the Laplace sharpness and five percent threshold constitute the result feedback, together forming a typical "prior-measurement-correction" closed-loop control framework; this framework enables the focal plane slip to adapt to external disturbances in real time without relying on manual adjustment, thereby ensuring that the contact network conductor always enters the subsequent multi-feature fusion and outlier detection process with near-optimal clarity in each local window of the mirrored grid sequence, laying a high-quality, quantifiable and continuously stable source data foundation for the contact network safety inspection and side-by-side inspection device based on image recognition.
[0039] Furthermore, when the improvement in the comprehensive sharpness score reaches a threshold and is determined to be effectively focused, the system immediately tags the current visible light image and infrared image with the same label and writes them as a pair of in-focus frames to the cache. Simultaneously, the grid index, focus timestamp, comprehensive sharpness score, and image distance increment are written to the access record table, providing the grid with the latest state snapshot. This serves two fundamental purposes: first, the cache maintains a data window closest to real time in a ring structure, providing a low-latency source for subsequent multi-feature fusion; second, the access record table retains grid-level time series, which can be used to evaluate the stability of focus adjustment strategies under different operating conditions. Because the mirrored grid sequence drifts overall every few frames, the same grid index may be accessed continuously within a short period of time. To prevent the cache from flooding with homogeneous data, the system performs a phase consistency comparison on any two in-focus frames with the same grid index and an access interval of less than 60 frames. The comparison algorithm calculates three indicators, namely grayscale mean square error, infrared radiation mean difference and edge gradient difference, and compares them with the empirical threshold. If all three indicators are lower than the threshold, it means that the two frames are almost identical in terms of brightness, thermal radiation and structural edge, which can be regarded as information redundancy. At this time, the system only retains the in-focus frames with newer timestamps, discards the older frames, and updates the latest timestamp and sharpness benchmark in the raster access record. Through this "write-comparison-cropping" closed loop, the intelligent analysis end ensures that the cache always contains the latest and most representative local imaging, while avoiding memory usage and subsequent feature statistical deviations caused by the accumulation of redundant data, thereby maintaining the effective density and processing efficiency of the data stream in high-speed operation scenarios.
[0040] Furthermore, after acquiring a pair of clearly focused frames, the intelligent analysis end immediately calculates nine complementary components around the same grid index. The first three components focus on visible light information: the mean visible light brightness measures the overall grayscale energy level, the mean visible light gradient directional distribution reflects the dominant direction of texture orientation, and the visible light edge density characterizes the abundance of structural lines. The middle three components are derived from the infrared spectrum: the mean infrared radiation intensity reveals the average energy level of the surface temperature field, the mean infrared radiation gradient directional distribution corresponds to the dominant direction of the thermal field distribution, and the infrared temperature difference amplitude quantifies the extreme difference between hot and cold in the same area. The next two components specifically describe cross-spectral correlation: the cross-spectral histogram measures the covariance of the visible light brightness and infrared radiation grayscale distribution within the grid, and the cross-spectral phase consistency evaluates the degree of synchronization of second-order structure edges in multiple spectral domains. The last component is the focal plane image distance increment, which explicitly incorporates the adjustment amplitude of this local focusing behavior into the feature vector, making the optical system state a learnable signal.
[0041] The nine components are sequentially concatenated to form a raster multi-component descriptor of length 9, and are uniformly stored in 32-bit floating-point format to ensure the numerical dynamic range and the accuracy of subsequent matrix operations. The system stipulates that no fewer than 60 raster multi-component descriptors must be collected for the same mirrored raster index within the first 120 frames after the train starts. The design of this "cold start" window is essentially to quickly capture the statistical profile of the raster under various instantaneous working conditions through short-term high-frequency sampling. After collecting a sufficient number of samples, the statistical module calculates the mean and standard deviation of the nine components and performs normalization accordingly, so that physical quantities of different dimensions are mapped to the same zero-mean, unit variance space. This not only weakens the direct impact of brightness fluctuations and temperature seasonal drift on the model, but also prevents any single component from dominating the learning process due to its excessively large numerical range. The normalized raster multi-component descriptors are fed into a three-layer fully connected network with a layer-by-layer contraction structure of 27, 18, and 9 nodes. The activation function uses rectified linear units with a leakage coefficient of 0.01 to balance nonlinear expressiveness and gradient flow stability. The network is designed to be equivalent to a lightweight multi-spectral fusion mapper: it projects the original nine-dimensional space back into nine dimensions after two nonlinear blending operations. However, the dimensions of the output vector are no longer simple components but high-order semantic factors that combine spectral, spatial, and optical states. The system concatenates this output vector with the raster index to form a fused vector, which is written to the candidate buffer in real time. The candidate buffer is based on a two-ended architecture, with new entries appended incrementally over time to maintain the latest state of the current observation window for subsequent reference vector matching and outlier detection. Through this complete process, the intelligent analysis end can compress local multimodal information into a uniformly scaled, measurable fused vector that incorporates the machine's own state within milliseconds, laying a highly consistent feature foundation for subsequent online background updates and potential anomaly flagging.
[0042] Furthermore, the intelligent analysis end considers each fused vector as a high-dimensional projection of the current grid state, while the reference vector library acts as the grid's "long-term memory." The distance between the two measures the degree of deviation between the instantaneous observation and the long-term baseline. The reference vector library is updated online with an exponential decay coefficient of 0.95. Its mathematical essence is a first-order recursive averaging of historical fused vectors, giving a slightly higher weight to recent data and an exponentially decreasing weight to older data. This allows the library vectors to slowly track natural drift due to seasonal light, temperature, and mechanical wear, while also being immune to rapid fluctuations due to transient noise. The distance between the current fused vector and the reference vector is calculated using the Euclidean norm, with a distance threshold set at 2.5 times the historical standard deviation of the reference vector. This is based on the empirical assumption that the fused vector approximates a Gaussian distribution in high-dimensional space, and a tail probability of approximately 0.012 is used to identify outliers. If the distance exceeds the threshold, it indicates a low-probability grid state. The fused vector is then registered as a potential anomaly and added to the defect candidate set, along with the grid index, timestamp, and distance value.
[0043] The defect candidate set uses a double-ended queue, adding new entries incrementally over time. The oldest entry is deleted when the set exceeds 1200. This sliding window design ensures that the system retains only recent data relevant for real-time analysis, preventing long-tail history from overwhelming current anomalies. Each time a potential anomaly enters the set, the inference and judgment process is triggered. The process first searches for entries within the set with a time interval of no more than 15 frames and a grid index difference of no more than 3, grouping them into the same physical cluster. This is based on the physical mechanism that the same defect will fall into adjacent grids and be continuously observed when the mirrored grid sequence is slightly shifted. The maximum distance value within the same cluster is used as the cluster distance value to ensure the most conservative upper bound for the anomaly metric. The global frame number of each cluster entry is then mapped to a segment identifier using a segment index table. If a cluster appears in two or more consecutive segments, it is marked as consistent across segments, indicating that the anomaly is stable and reproducible over time, rather than being a sporadic noise. The system then takes a weighted average of the cluster distance values and the number of segments that appear for the cross-segment consistent clusters to obtain the confidence level. The weight design allows the distance value to reflect the intensity of the anomaly, and the number of segments to reflect the persistence of the anomaly. If the confidence level is higher than 0.8, the cluster entry will be included in the suspected defect set. To prevent the same physical defect from being recorded repeatedly under grid drift and segment division, the system uses a union-find algorithm to merge entries in the suspected defect set with a grid index difference of no more than 2 and the same segment identifier. Only the one with the highest confidence level is retained as the representative entry, and the grid index, segment identifier, confidence level, and first frame timestamp are recorded. The core principle of the entire link is to gradually reduce the statistical threshold for the definition of "anomaly": first, outliers are captured with high-dimensional distances, and then physical clusters are formed using time and space proximity constraints. Then, random noise is filtered using cross-segment consistency and confidence levels, and finally, duplications are eliminated by merging with the help of a union-find algorithm. This not only ensures the detection sensitivity, but also has a strong suppression capability against environmental disturbances and occasional errors, so that the overhead line safety inspection and patrol device based on image recognition can calibrate high-confidence suspected defects in real time and robustly in the continuously changing train operation environment.
[0044] The following example lists the maximum operating speeds The high-speed EMU is used as the subject to fully demonstrate the whole chain implementation process of the overhead line safety inspection device based on image recognition from outbound calibration to manual review. Perform multi-step micro-shift search on the optical axis to measure the zero focal length of the visible light camera and infrared camera respectively 、 The optical axes of the two machines are co-pointed and the optical center distance is , this value falls within the design constraints Phase-locked loop reference frequency Generation time stamp resolution , any frame timestamps in the following text are An integer multiple of .
[0045] After the train starts, the acquisition end uses the frame rate 、 Start synchronous acquisition resolution Visible light images and resolution Infrared image. Frame synchronization error ,in is the time stamp of the visible light and infrared frames. Trigger threshold, control firmware to multiply visible light frame rate to , and start the four-level pyramid downsampling (scaling factor ), perform row and column prediction compression on adjacent level differences, and the instantaneous code stream after compression ,satisfy .
[0046] The mirror grid plane radius is set to . Golden Angle Radial spiral layout seed points, Point polar coordinates The dot spacing is maintained at . Perform Thiessen segmentation to obtain a honeycomb grid with an average area , area variance .
[0047] Every Frame grid overall rotation and translates along the direction of train movement . Define the coverage function , system real-time monitoring to ensure For the same time frame, the intelligent analysis end maps the pixels to the current grid according to the camera's internal and external parameters, using sub-pixel bilinear interpolation. , and uniquely belongs to the grid .
[0048] Iterate over the list from the grid index Increasing from Then backtrack to form a round-trip loop. Read the last image distance increment , combined with speed , vibration acceleration , ambient illumination Interpolation ,Pick , , , , and limit . Drive the focal plane to move Post-stationary Calculate the Laplace energy for visible light and infrared images separately , comprehensive sharpness .
[0049] like , it is recorded as a clear focus frame and cached, and recorded at the same time . For the same grid and interval Calculate the grayscale mean square error of two pairs of clear frames , infrared radiation mean difference , edge gradient difference ,like , then the old frame is discarded. In the reserved frame, the grid Calculate nine-dimensional components Before startup Frame Collection Samples, find the mean , standard deviation , normalized .
[0050] The normalized vector is passed through a three-layer fully connected network (node , leakage coefficient ) generates output . Reference vector library to raster maintain , online update .
[0051] distance .
[0052] like Register as a potential exception and add it to the double-ended queue , capacity upper limit Each queue triggers clustering: if the entry time difference Frame and grid difference Cluster , cluster distance .
[0053] The global frame number to segment identifier mapping gives the number of segments that appear Confidence .
[0054] like Enter suspected defect set .exist Same segment and grid difference Execute the union-lookup algorithm and keep the entry with the highest confidence The entry is pushed to the observation tablet via a link, which renders a multimodal thumbnail and allows manual annotation. They are real-time speed, vibration acceleration, and illumination; is the spatial focal length increment; are the mean brightness and radiation respectively; is the edge density; is the cross-spectral histogram; is cross-spectral phase consistency; is the maximum distance of the cluster; is the cluster confidence.
[0055] Figure 2 This paper presents the performance comparison results of a mirror-based progressive multi-dimensional fusion defect inference algorithm in an image recognition-based overhead line safety inspection device. This experiment compared the performance of the proposed progressive depth-of-focus strategy with a fixed focal length approach in terms of comprehensive sharpness scores under different train operating conditions. In the experimental setup, the horizontal axis represents train speed, ranging from 60 km / h to 360 km / h, covering typical high-speed train operating conditions. The vertical axis represents the comprehensive sharpness score calculated via Laplace transform, which reflects the focus clarity of the visible and infrared images within the virtual mirror plane. The experimental results show that the proposed device (solid curve), employing the progressive depth-of-focus strategy, maintains a high comprehensive sharpness score across the entire speed range. At a train speed of 60 km / h, the comprehensive sharpness score reaches 0.82; even at a high-speed of 360 km / h, the score remains above 0.77. This is due to the motorized focus mechanism in the present invention, which calculates the image distance increment in real time based on the grid index of the current mirrored grid sequence, combined with parameters such as train speed, vibration amplitude, and ambient light intensity, and drives the focal plane to precisely slide within a ±0.20 mm range. In contrast, the conventional device using a fixed focal length focusing method (dashed curve) shows a significant decrease in overall sharpness score as train speed increases. The score is 0.72 at 60 km / h, but drops to 0.32 at 360 km / h, failing to meet the clarity requirements for high-speed inspections. Experiments also verified the effectiveness of the cross-segment consistency comparison mechanism. When the confidence level is greater than 0.8 and occurs in two or more consecutive segments, the system can accurately identify suspected defect sets. Final test results show that the present device achieved a 95.3% defect detection rate, a false alarm rate of only 2.1%, and processing latency within 1 millisecond, fully demonstrating the technical advantages of the mirrored progressive multi-dimensional fusion defect inference algorithm for contact line safety inspection applications.
[0056] Figure 3The working principle of the progressive depth-of-focus strategy in the overhead catenary safety inspection device is demonstrated. This strategy is a core component of the mirror-based progressive multi-element fusion defect inference algorithm, used to achieve precise focused imaging of the overhead catenary during train operation. The upper portion of the figure depicts the physical structure of the overhead catenary, including the horizontal main line and three vertical pillars. As the power supply for electric locomotives, the overhead catenary is the primary inspection target of the device. The area between the pillars constitutes the target area for defect detection. The core component of the device is the dual-camera system located in the lower portion of the figure. This system consists of a visible light camera and an infrared camera, coaxially mounted and sharing a common optical axis, with an optical center distance of less than 30 mm. The dual-camera system uses a motorized focus mechanism to achieve precise adjustment of the focal plane. The working mechanism of the progressive depth-of-focus is demonstrated in the three elliptical focal plane positions at different heights in the figure. Focal plane positions 1, 2, and 3 represent three consecutive focus states during the traversal of the mirror grid sequence. The straight line connecting the centers of each focal plane circle illustrates the path of focal plane sliding, embodying the technical feature of "grid-by-grid focal plane sliding." In practice, the intelligent analysis end traverses the grid in ascending order, starting with the grid with the smallest index, based on the grid index of the current mirrored grid sequence. For each target grid, the system reads the image distance increment recorded during the previous visit and, combining parameters such as the train's current speed, vibration amplitude, and ambient light intensity, calculates the new image distance increment in real time using a linear interpolation table. This increment, serving as the focal plane sliding step size, is strictly limited to a range of ±0.20 mm along the optical axis to ensure focus accuracy. When traversing into a target grid, the motorized focus mechanism drives the visible light camera and infrared camera to move the focal plane uniformly and at a constant speed along the optical axis. After the focal plane moves, it remains stationary for at least 0.50 milliseconds to eliminate the coupling effects of mechanical inertia and vehicle body vibration. During this stationary period, the system uses Laplace transforms to calculate the full-frame sharpness scores for the current visible light and infrared images, respectively, and calculates a weighted average to obtain a composite sharpness score. When the comprehensive sharpness score is improved by at least 5% compared to the previous visit, the current focus is determined to be valid, the current frame is marked as a focused frame and written to the cache.
[0057] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these specific embodiments are merely illustrative, and that those skilled in the art may omit, substitute, and modify the details of the methods and systems described above without departing from the principles and spirit of the present invention. For example, combining the above method steps to perform substantially the same functions in substantially the same manner to achieve substantially the same results falls within the scope of the present invention. Accordingly, the scope of the present invention is limited solely by the appended claims.
Claims
1. A device for conducting safety inspection of overhead line based on image recognition, characterized in that: The device includes: an acquisition end, an intelligent analysis end and an observation tablet; the acquisition end is fixedly installed on the train, and synchronously acquires continuous visible light images and infrared images of the contact network during the operation of the train, and constructs an original multimodal data stream queue in chronological order in the local cache, divides the original multimodal data stream queue into equal-length data segments according to preset frame group rules, and transmits them to the intelligent analysis end in real time via a high-speed, low-latency link; the intelligent analysis end calls the mirror progressive multi-fusion defect inference algorithm, and performs the following specific processes on each data segment: generates a mirror grid sequence, and grid-matches the visible light image and the infrared image in the virtual mirror plane; adopts a progressive depth focus strategy, realizes focal plane sliding grid by grid in the mirror grid sequence, and extracts a clear-focus frame after each sliding; performs multi-feature fusion on the clear-focus frame, and the fusion result enters the defect candidate set; executes the inference judgment process on the defect candidate set, and the inference judgment process compares the candidate set across segments according to the segment index table, and outputs the suspected defect set and the corresponding segment identifier; synchronizes the suspected defect set with the corresponding segment identifier The data is pushed to the observation tablet, and the inspection personnel manually review the suspected defect set through the observation tablet; the intelligent analysis end calculates the following components for the visible light image and infrared image with the same grid index in each pair of clear focus frames: visible light brightness mean, visible light gradient direction distribution mean, visible light edge density, infrared radiation intensity mean, infrared radiation gradient direction distribution mean, infrared temperature difference amplitude, cross-spectral histogram, cross-spectral phase consistency and focal plane image distance increment; the above 9 components are spliced in sequence to obtain a grid of length 9 Multi-component descriptors are stored in 32-bit floating-point format. For the same mirrored raster index, no less than 60 raster multi-component descriptors are collected within the first 120 frames after the train starts. The mean and standard deviation of each component are calculated and normalized. The normalized raster multi-component descriptors are input into a 3-layer fully connected network. The number of nodes in each layer of the network is 27, 18, and 9, respectively. The activation function uses a rectified linear unit with a leakage coefficient of 0.
01. The output vector and the raster index together constitute a fusion vector, which is written into the candidate buffer queue in real time.
2. The overhead line safety inspection and patrol device based on image recognition according to claim 1, characterized in that: The acquisition end is equipped with a visible light camera and an infrared camera, which are coaxially installed and share the optical axis center. The optical center distance between the two does not exceed 30 mm; the single-frame spatial resolution of the visible light camera is not less than 1920×1080 pixels, and the frame rate is not less than 60 frames per second; the detection band of the infrared camera is 8 microns to 14 microns, the single-frame spatial resolution is not less than 640×512 pixels, and the frame rate is not less than 60 frames per second.
3. The overhead line safety inspection and patrol device based on image recognition according to claim 2, characterized in that: During the acquisition process of visible light images and infrared images, the acquisition end writes a unified time stamp for each frame, and uses a phase-locked loop and train speed pulse signal for synchronous calibration to keep the inter-frame time difference between the two types of images within 1 millisecond; when the train speed exceeds 200 kilometers per hour, the acquisition end automatically increases the visible light camera frame rate to 120 frames per second and enables spatial four-level downsampling differential compression to ensure that the actual bandwidth occupied by the link does not exceed 80% of the rated bandwidth.
4. The overhead line safety inspection and patrol device based on image recognition as claimed in claim 3, characterized in that: The process of generating a mirror grid sequence on the intelligent analysis end includes: taking the intersection of the optical axes of the two cameras as the plane origin, laying out seed points in the virtual mirror plane according to the radial spiral method with the golden angle as the increment, with the number of seed points being no less than 800, and the plane distance between any two adjacent seed points being maintained between 6 mm and 8 mm; using the Thiessen polygon segmentation method based on the seed points to form a honeycomb mirror grid sequence, and controlling the area variance of a single grid in the mirror grid sequence to within 5% to ensure that the area of each grid is approximately consistent; during the operation of the train, all seed points are rotated 137.5 degrees according to the golden angle and translated by 1 pixel along the direction of the train's movement every 5 frames, and a new mirror grid sequence is generated in real time, thus forming a time-progressive mirror grid sequence.
5. The overhead line safety inspection and patrol device based on image recognition as claimed in claim 4, characterized in that: The mirror grid sequence ensures complete coverage of the virtual mirror plane within any continuous 60 frames; the intelligent analysis end performs perspective mapping on the visible light image and infrared image acquired at the same time to the current mirror grid sequence. Sub-pixel precision bilinear interpolation is used in the mapping process to uniquely attribute each pixel to the corresponding grid.
6. The overhead line safety inspection and patrol device based on image recognition according to claim 5, characterized in that: Before the train departs, the acquisition end performs an autofocus process on both the visible light camera and the infrared camera, determines the optical zero-point focal length of the train when stationary, and writes this focal length into the control register as the focal length reference for the entire process. The intelligent analysis end traverses the list grid by grid in ascending order, starting with the grid with the smallest index based on the grid index of the current mirrored grid sequence. When traversing to the grid with the largest index, it immediately backtracks in the opposite direction, forming a round-trip loop sequence, ensuring that each grid is visited twice within a continuous traversal cycle with a balanced visit interval. For each grid in the traversal list, the image distance increment value recorded during the previous visit to the grid is read. Combined with the current train speed, vibration amplitude, and ambient light brightness, the linear interpolation table is called to calculate the new image distance increment value in real time. The obtained new image distance increment value is used as the step size of this focal plane sliding, and the step size value range is limited to within ±0.20 mm in the optical axis direction. When the traversal process enters a certain target grid, the electric focus mechanism of the visible light camera and the infrared camera is driven to move the focal plane along the optical axis at a uniform constant speed, and the moving distance is equal to the newly added image distance increment value; the focal plane remains stationary for at least 0.50 milliseconds after movement to eliminate the coupling effect of mechanical inertia and vehicle body vibration; during the time slot when the focal plane completes the sliding and remains stationary, the Laplace transform is called to calculate the full-frame sharpness score of the current visible light image and infrared image respectively; the full-frame sharpness scores of the two images are weighted averaged, and the resulting comprehensive sharpness score is compared with the comprehensive sharpness score when the grid was last visited. If the improvement is at least 5%, the current focus is determined to be valid.
7. The overhead line safety inspection and patrol device based on image recognition according to claim 6, characterized in that: When the intelligent analysis end determines that the current focus is valid, it marks the current visible light image and infrared image as in-focus frames and writes them to the cache; at the same time, it records the current grid index, focus timestamp, comprehensive sharpness score, and image distance increment value, and updates the access record of the corresponding grid; for any two in-focus frames with the same grid index and an access time interval of less than 60 frames in the same mirrored grid sequence, a phase consistency comparison is performed; if the comparison results are lower than the threshold in the three indicators of grayscale mean square error, infrared radiation mean difference, and edge gradient difference, the in-focus frame with a newer timestamp is retained and the older one is discarded to reduce redundant data.
8. The overhead line safety inspection and patrol device based on image recognition according to claim 7, characterized in that: For each fused vector, the intelligent analysis client queries the reference vector library for the corresponding raster index. The reference vector library is updated online with an exponential decay coefficient of 0.
95. The distance between the current fused vector and the reference vector is measured, with the distance threshold preset to 2.5 times the historical standard deviation of the reference vector. If the distance exceeds the threshold, the fused vector is registered as a potential anomaly and written into the defect candidate set along with the grid index, timestamp, and distance value. The defect candidate set uses a double-ended queue structure, with new entries added in chronological order. Once the set size exceeds 1200 entries, the oldest entry is removed based on the first-in, first-out principle. Each new potential anomaly triggers the inference and judgment process immediately. The inference judgment process includes: entries in the defect candidate set with a time interval of no more than 15 frames and a grid index difference of no more than 3 are regarded as the same physical cluster; the maximum distance value within the same cluster is taken as the cluster distance value; according to the segment index table, the global frame number of each cluster entry is mapped to the segment identifier; if the same cluster appears in 2 or more consecutive segments, the cluster is marked as cross-segment consistent; the cluster distance value of the cross-segment consistent cluster and the number of segments where it appears are weighted averaged to obtain the confidence level; when the confidence level is higher than 0.8, the cluster entry enters the suspected defect set; for entries in the suspected defect set with a raster index difference of no more than 2 and the same segment identifier, a union-find algorithm is used to perform a merge operation, and the entry with the highest confidence level is retained as the representative entry; each representative entry contains the raster index, segment identifier, confidence level and first frame timestamp.