Real-time railway 2C analysis system and method based on deep learning
Through deep learning technology, combined with multimodal data processing of visible light and infrared images, high-precision dynamic monitoring of the contact line is achieved, which solves the real-time and accuracy issues of contact network status detection, provides stable assessment and risk warning of contact line geometric changes, and improves the safety and maintenance efficiency of the railway power supply system.
Patent Information
- Application Number
- CN202511095771.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing contact network status detection technology suffers from insufficient accuracy and poor real-time performance in high-speed and heavy-load railways. In particular, it is difficult to accurately monitor the geometric state of the contact line under complex lighting and background interference conditions, affecting the train's power quality and driving safety.
A real-time railway 2C analysis system based on deep learning is adopted. Visible light and infrared image frames are synchronously acquired through a multimodal data acquisition unit. Combined with a multi-dimensional recursive coupling analysis unit and a composite risk index calculation unit, high-precision quantification of the vertical and horizontal deviations of the contact line is achieved. A fused pseudo-color tensor is generated and multi-scale convolution kernel group feature extraction is performed. The center curve of the contact line is determined using iterative morphological skeletonization and guide curve tracking, and the composite risk index is calculated.
It significantly improves the intelligence and reliability of contact network status monitoring, can dynamically capture contact line geometric changes under complex lighting and background conditions, provide high-precision deviation assessment and real-time risk warning, and improve the operational safety and maintenance efficiency of the railway power supply system.
Smart Images

Figure CN120598947B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning technology, and specifically relates to a real-time railway 2C analysis system and method based on deep learning. Background Art
[0002] In rail transit power supply systems, the catenary (also known as the contact wire or contact conductor) is a critical structure for power transmission. The stability of its geometric state is directly related to the quality of train power supply and operating safety. With the rapid development of high-speed and heavy-haul railways, the catenary system faces more stringent operating conditions. Factors such as increased train speeds, increased traction currents, and complex environmental disturbances have placed higher requirements on the accuracy and real-time performance of catenary state detection technology. Therefore, how to achieve real-time monitoring of the geometric state of the catenary during train operation has become a key research topic in the field of rail transit safety.
[0003] Currently, the industry has proposed a variety of contact wire detection and monitoring technologies, including laser ranging, image processing, infrared thermal imaging, and multi-sensor fusion. Laser ranging uses a laser and reflector mounted on the roof of the train to achieve high-precision measurement of contact wire height and pullout. While this method offers high accuracy in static testing, it also faces challenges such as high equipment costs, difficulty in on-site deployment, and sensitivity to ambient light interference. Furthermore, laser ranging typically relies on reflected light intensity, and the contact wire surface material or contamination can lead to uncontrollable ranging errors, making long-term stable operation difficult.
[0004] Image processing methods use high-speed industrial cameras to capture images of the contact line, extract its contours through image recognition algorithms, and analyze its positional changes. This method is low-cost and suitable for installation on the roof of high-speed trains for dynamic detection, making it a widely used technical approach. However, most existing image processing methods rely on single-modal images, processing only visible light images. This makes them susceptible to factors such as lighting changes, background interference, and image blur. For example, in low-quality imaging scenarios such as tunnels, at night, or in strong backlight, the contrast between the contact line and the background decreases significantly, making it difficult for image algorithms to accurately segment the target, resulting in distorted detection results. Summary of the Invention
[0005] In view of this, the main purpose of the present invention is to provide a real-time railway 2C analysis system and method based on deep learning, which significantly improves the intelligence and reliability of contact network status monitoring.
[0006] The technical solution adopted in the present invention is as follows:
[0007] A real-time railway 2C analysis system based on deep learning includes: a multimodal data acquisition unit, a multidimensional recursive coupling analysis unit, and a composite risk index calculation unit; the multimodal data acquisition unit is used to synchronously acquire visible light image frames and infrared image frames at a fixed frame rate during train operation using image acquisition devices configured along the line, synchronously record positioning coordinates and timestamps, and align the timestamps of the visible light image frames and infrared image frames to obtain a 2C feature frame sequence; the multidimensional recursive coupling analysis unit is used to construct a continuous frame stack of the 2C feature frame sequence according to a predetermined length and slide it with a set step size to regenerate a fused pseudo-color tensor; a multi-scale convolution kernel group is used to extract gradient and phase features step by step; the center curve of the contact line is determined by iterative morphological skeletonization and guide curve tracking to obtain its vertical deviation and horizontal deviation; the composite risk index calculation unit is used to calculate the composite risk index of the contact network at the corresponding moment based on the vertical deviation and horizontal deviation.
[0008] Furthermore, the multimodal data acquisition unit aligns all visible light image frames and infrared image frames with the same timestamp and uniformly resamples them to a resolution of 1mm:1px to obtain a normalized 2C feature frame sequence. ; Indicates time; the bit depth of each visible light image frame or infrared image frame is 8 bits per pixel.
[0009] Furthermore, the multi-dimensional recursive coupling analysis unit will By length Build a stack of consecutive frames and step Sliding; performing linear brightness stretching on each visible light image frame, raising the darkest pixel grayscale to 0 and compressing the brightest pixel grayscale to 255; performing temperature range remapping on each infrared image frame, mapping the lowest radiation value to 0 and the highest radiation value to 255; performing homography transformation on the visible light image frame and the infrared image frame, aligning them pixel by pixel into a completely overlapping two-dimensional alignment grid, and confirming that the visible light pixel and the infrared pixel at any row and column index position point to the same spatial point.
[0010] Furthermore, the multimodal data acquisition unit performs the following operations on each pixel in the aligned grid: the brightness value of the visible light pixel is directly used as the first color component; the grayscale value of the infrared pixel is mapped to the same 0 to 255 range through linear interpolation as the second color component; the third color component is constructed with the average value of the first color component and the second color component to enhance the overall contrast between the contact line and the background; the first color component, the second color component and the third color component are spliced at the same row and column index position at the channel level according to the row priority order to generate a fused pseudo-color tensor whose height, width and number of channels correspond to the number of pixel rows, the number of pixel columns and 3 respectively.
[0011] Furthermore, when the multi-dimensional recursively coupled parsing unit uses a multi-scale convolution kernel group to extract gradient and phase features step by step, the multi-scale convolution kernel group includes three preset groups of convolution kernels with kernel sizes of 3×3, 7×7 and 11×11, respectively.
[0012] Furthermore, each convolution kernel group contains a horizontal direction operator and a vertical direction operator; the convolution kernel group of size 3×3 is slid on the fused pseudo-color tensor at a pixel-by-pixel step size, and the local response obtained by the horizontal difference and the vertical difference is calculated in each channel; the first layer gradient amplitude map and the first layer phase angle map are generated according to the relative size of the horizontal difference and the vertical difference; the convolution kernel group of size 7×7 is slid on the fused pseudo-color tensor at a pixel-by-pixel step size, and the first layer gradient amplitude map and the first layer phase angle map are reconvolved; the directional difference is recalculated on the same channel to obtain the second layer gradient amplitude map and the second layer phase angle map; the convolution kernel group of size 11× The convolution kernel group of 11 slides on the fused pseudo-color tensor with a pixel-by-pixel step size, and performs convolution operation on the second-layer gradient amplitude map and the second-layer phase angle map; outputs the third-layer gradient amplitude map and the third-layer phase angle map; according to the pixel correspondence, the three-layer gradient amplitude map is compared for the maximum value at the same position, and the maximum gradient amplitude is retained. At the same time, the arithmetic average of the three-layer phase angle map is taken to eliminate isolated directional fluctuations; the global gradient amplitude map and the global phase angle map are obtained in the same output space; according to the row priority order, the global gradient amplitude map and the global phase angle map are spliced in the channel dimension to generate a gradient phase matrix with the size of the number of pixel rows, the number of pixel columns and 2.
[0013] Furthermore, the multi-dimensional recursively coupled parsing unit uses a fixed threshold method to determine whether the gradient amplitude value of each pixel exceeds a set threshold in the global gradient amplitude map, marks pixels exceeding the threshold as foreground, and marks pixels below or equal to the threshold as background, generating an initial binary map containing only two types of values: foreground and background; subsequently, an opening operation and a closing operation are performed on the initial binary map in sequence to remove isolated foreground areas with an area less than 4 pixels, and then a skeleton map is obtained through iterative morphological skeletonization, including: with the pixel as the center, eliminating areas with an area less than 4 pixels in a three-by-three neighborhood according to the set structural element order; after completing one traversal, it is determined whether there are still deletable areas in the initial binary map; if so, the next traversal is continued; when no change occurs in two consecutive traversals, the skeletonization process converges to obtain a skeleton map with a single pixel width.
[0014] Furthermore, the process of obtaining the vertical deviation and horizontal deviation of the multi-dimensional recursive coupling parsing unit includes: scanning the skeleton graph, identifying pixels with a degree of 1 as endpoints, retaining the longest branch associated with the endpoint, and deleting all other branches with a length of less than 15 pixels; among the skeleton endpoints, selecting the one with the smallest row coordinate as the starting point; based on the four-connectivity rule, writing its adjacent skeleton pixels in the order of column coordinates from small to large into the trajectory queue as the initial tracking path; repeating the following operations until reaching the other endpoint: reading the current pixel from the end of the trajectory queue, searching for unvisited pixels adjacent to the current pixel in its neighborhood, and if there are multiple candidates, giving priority to the candidate. Select pixels with larger column coordinates, add the selected pixels to the trajectory queue and mark them as visited; calculate the center point of the complete trajectory queue using a five-point sliding window, and connect the centers of consecutive windows into a smooth broken line; for positions where the angle change in the broken line exceeds 30 degrees, use the three-point averaging method to eliminate sharp turns and obtain the final contact line center curve; using the train coordinate system as a reference, project the final contact line center curve onto the vertical and horizontal axes according to row and column coordinates: calculate the average value of the row coordinates of all pixels on the final contact line center curve, and the difference between it and the design elevation is the vertical deviation; calculate the average value of the column coordinates of all pixels on the center curve, and the difference between it and the design center line is the horizontal deviation.
[0015] Furthermore, the composite risk index calculated by the composite risk index calculation unit is:
[0016] ;
[0017] in, For time Composite risk index at the time of For the moment Vertical deviation, unit: mm; For the moment Horizontal deviation, unit: mm; is the vertical tolerance, unit is mm; is the horizontal tolerance, in mm; is the instantaneous change of vertical deviation, unit is mm; It is the instantaneous change of horizontal deviation, in mm.
[0018] A real-time railway 2C analysis method based on deep learning is disclosed. The method comprises: during train operation, image acquisition devices configured along the line synchronously acquire visible light image frames and infrared image frames at a fixed frame rate, synchronously record positioning coordinates and timestamps, and align the timestamps of the visible light image frames and infrared image frames to obtain a 2C feature frame sequence; constructing a continuous frame stack of the 2C feature frame sequence according to a predetermined length and sliding it with a set step size to regenerate a fused pseudo-color tensor; using a multi-scale convolution kernel group to extract gradient and phase features step by step; determining the center curve of the contact line through iterative morphological skeletonization and guide curve tracking to obtain its vertical deviation and horizontal deviation; and calculating the composite risk index of the contact network at the corresponding moment based on the vertical deviation and horizontal deviation.
[0019] The above technical solution achieves the following beneficial effects: It fully integrates visible light and infrared image modal information to construct a uniformly aligned, uniformly resolved, and channel-fused 2C feature frame sequence, effectively enhancing image representation under complex lighting and background conditions. By introducing a multi-dimensional recursive coupled parsing mechanism, it achieves temporal modeling and spatial feature extraction for consecutive image frames. Combined with a multi-scale convolutional structure, it progressively captures local and global edge information of the contact line, enhancing the response consistency and recognition robustness of the contact line under different operating conditions. Furthermore, the system constructs a skeleton structure by fusing gradient and phase images and extracts the contact line center curve using a guided curve tracing algorithm, resulting in a stable and continuous structural trajectory. Compared to traditional single-frame static recognition methods, the present invention dynamically captures the geometric trajectory of the contact line and, combined with trajectory center point calculation, achieves high-precision quantification of vertical and horizontal deviations, ensuring structural continuity and temporal consistency in deviation assessment. Furthermore, the system incorporates a composite risk index calculation mechanism to comprehensively analyze deviation amplitude and variation trend, enabling quantitative risk assessment and trend determination of the contact network status, providing strong early warning capabilities and real-time performance. In general, the present invention has the advantages of strong data fusion capability, high spatial modeling accuracy, reasonable risk quantification, and strong system adaptability. It can be widely used in the intelligent operation and maintenance system of contact networks in the fields of high-speed railways, heavy-load railways and urban rail transit, and can significantly improve the operational safety and maintenance efficiency of railway power supply systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic diagram of the system structure of a deep learning-based real-time railway 2C analysis system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] All features disclosed in this specification, or all steps in the disclosed methods or processes, except mutually exclusive features and / or steps, can be combined in any manner.
[0022] Any feature disclosed in this specification (including any appended claims and abstract), unless otherwise stated, may be replaced by other equivalent or similar features. In other words, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.
[0023] refer to Figure 1 , a real-time railway 2C analysis system based on deep learning, the system includes: multimodal data acquisition unit, multi-dimensional recursive coupling analysis unit and composite risk index calculation unit;
[0024] The multimodal data acquisition unit is a key component of the deep learning-based real-time railway 2C analysis system. Its core task is to acquire and process image data in two distinct modalities: visible light and infrared, providing the fundamental input for subsequent analysis. During train operation, image acquisition devices installed along the track continuously capture the overhead contact network at a fixed frame rate. Each acquisition device is equipped with both a visible light camera and an infrared camera, enabling simultaneous capture of visible and infrared image frames at corresponding spatial locations at any given moment. Because these two modalities reflect distinct characteristics of the target's physical environment—visible light images provide rich texture and structural details, while infrared images are sensitive to the target's thermal radiation characteristics—combining the two provides more comprehensive observational information, improving the reliability and accuracy of overhead contact network condition analysis. Each image frame is accompanied by the corresponding high-precision positioning coordinates and accurate timestamp. A high-precision clock synchronization mechanism ensures that data collected by all sensors at the same moment is strictly synchronized. Subsequently, by accurately matching and aligning the recorded timestamps, the visible light image frames and the infrared image frames are matched one-to-one, so that the two are completely synchronized in the time dimension, forming a 2C feature frame sequence with unified structure and time alignment.
[0025] The resulting 2C feature frame sequence is then further normalized. Specifically, each frame undergoes a unified resampling process to a standard spatial resolution of 1 mm to 1 pixel, eliminating resolution inconsistencies caused by differences in camera field of view or fluctuations in acquisition distance. Furthermore, to facilitate unified data processing, the bit depth of both visible and infrared images is standardized to 8 bits per pixel, thus avoiding the data processing complexity associated with different bit depths. A homography transformation is then performed in image space, using a precisely calibrated transformation matrix to precisely align the visible and infrared images pixel by pixel, ensuring that every pixel in the two modal images corresponds to the same physical point in space. Finally, at the image data level, these aligned images are further fused. Specifically, the visible light brightness of each pixel is used as the primary color channel, the infrared image grayscale value is linearly interpolated as another color channel, and a third color channel is constructed by averaging the two to enhance the contrast between the contact network target and the background. After this series of meticulous acquisition, synchronization, resampling and spatial alignment fusion processes, the fused pseudo-color image tensor output by the multimodal data acquisition unit is not only spatially and temporally consistent, but also possesses the combined advantages of visible light and infrared modalities, providing a reliable and high-quality data source for the subsequent multidimensional recursive coupling analysis unit.
[0026] The multi-dimensional recursive coupled parsing unit is a key component of the deep learning-based real-time railway 2C analysis system. It performs in-depth analysis of the fused pseudo-color tensor acquired by the front-end to extract crucial information representing the catenary's condition. First, the fused pseudo-color tensor, delivered from the multimodal data acquisition unit, is constructed into a continuous frame stack of a fixed length. This stack is then sequentially slid with a fixed step size to fully utilize data in both spatial and temporal dimensions, thereby capturing short-term trends in the catenary's condition. To enhance sensitivity to features in the image data, pixels within each frame undergo rigorous preprocessing, including linear brightness stretching for visible light frames and temperature remapping for infrared frames. A homography transformation is then used to ensure strict spatial correspondence between visible and infrared pixels, ensuring that the fused data accurately captures the comprehensive characteristics of the contact line and its surroundings. After this spatial alignment and brightness normalization, the multi-dimensional recursive coupled parsing unit extracts features from the fused pseudo-color tensor using multi-scale convolutional kernels of 3×3, 7×7, and 11×11 sizes, analyzing the data step-by-step. Convolution kernels of different scales play different roles. Smaller-sized convolution kernels can effectively extract detail features and reflect small local changes in the contact line, while larger-sized convolution kernels can capture structural features on a larger scale to enhance the expression of overall contour features.
[0027] Each set of convolution kernels includes directional operators for calculating horizontal and vertical differences. Through a step-by-step multi-scale convolution process, gradient magnitude and phase directional features are obtained at different locations in the fused image, accurately capturing subtle vertical and horizontal deviations of the contact line. After multi-scale convolution, the resulting gradient and phase feature maps are integrated into unified global gradient magnitude and phase angle maps to highlight the differences between the contact line target and background regions. Subsequently, a multi-dimensional recursively coupled parsing unit uses a fixed threshold method to identify foreground regions in the global gradient magnitude map, defining regions with significant gradient changes as foreground, thus forming an initial binary image. To further optimize the accuracy of this binary image, morphological opening and closing operations are used to eliminate noisy regions and isolated points, resulting in a clearer and more continuous contact line structure. These regions are then reduced to a single-pixel skeleton map using an iterative morphological skeletonization method to clearly define the geometric center of the contact line. Finally, a guided curve tracing method is used to select appropriate starting points within the skeleton map for pixel-by-pixel tracing, resulting in a smooth contact line center curve. By projecting this central curve, we can determine the vertical and horizontal deviations of the contact line relative to its design standard position. This multi-scale feature extraction and iterative skeleton optimization method gives the multi-dimensional recursively coupled analytical unit the advantages of stability and high precision, enabling accurate assessment of the operating status of the contact line.
[0028] The composite risk index calculation unit is a key functional module in the deep learning-based real-time railway 2C analysis system. It is responsible for further quantitative analysis of the contact line deviation information extracted by the multi-dimensional recursive coupling analysis unit to achieve real-time risk assessment of the catenary system. Specifically, the multi-dimensional recursive coupling analysis unit, through a series of image feature extraction, skeleton tracking, and smoothing processes, ultimately obtains the specific vertical and horizontal deviations of the contact line relative to the design standard position. The composite risk index calculation unit uses this as input to perform a comprehensive risk assessment of the catenary system. The basic principle of the composite risk index calculation unit is to comprehensively consider the magnitude and instantaneous trend of the vertical and horizontal deviations of the contact line at the current moment, thereby dynamically assessing potential safety hazards or abnormal conditions during catenary operation. In its implementation, the vertical and horizontal deviations are first compared with the system's pre-set allowable tolerance values, and the normalized deviation ratios are calculated. This normalization operation effectively eliminates the influence caused by the different scales of absolute values, ensuring that vertical and horizontal deviation information contribute equally to the risk index. Secondly, after obtaining the normalized values of the vertical deviation and the horizontal deviation, the instantaneous change is further considered, that is, the degree of change of the deviation in the vertical and horizontal directions at the current moment relative to the previous moment is calculated respectively, and these instantaneous changes are also normalized.
[0029] This approach aims to capture the dynamic characteristics of the contact line's state and promptly detect increases or decreases in transient risk. Next, the composite risk index calculation unit simultaneously considers the normalized information from both dimensions, using a specific mathematical combination to appropriately balance the weights of deviation magnitude and transient change to form a unified risk index. This calculation method considers the relative contribution of different deviation directions to the overall operational risk, thereby providing a comprehensive representation of the contact line's safety status. Furthermore, during the real-time calculation process, the composite risk index calculation unit continuously updates the composite risk index value at each point in time. This ensures the system maintains real-time awareness and rapid response to the current contact line state. Once deviations or transient changes exceed the permitted range, an early warning mechanism is triggered, providing decision support for railway operations management. Furthermore, to further enhance the stability of the system's risk prediction, the composite risk index calculation unit assigns higher weights to instances with larger deviation magnitudes during the numerical calculation process, emphasizing the sensitivity of the risk index and significantly increasing its sensitivity when significant anomalies occur, helping management personnel quickly detect and respond to potential risk events. Through this refined processing and effective integration of the relationship between vertical and horizontal deviations, instantaneous change trends and risk indices, the composite risk index calculation unit can provide accurate, stable and reliable risk assessment functions for the real-time railway 2C analysis system.
[0030] Furthermore, the core principle of the multimodal data acquisition unit for unified processing of visible light image frames and infrared image frames is to eliminate the differences in spatial position and time between data collected by different sensors through precise time synchronization and spatial resampling technology, so as to achieve complete consistency of the two image modalities. This process specifically includes frame-by-frame matching of visible light and infrared modal image frames with the same timestamp, thereby ensuring that the visible light image frames and infrared image frames acquired at each time point are strictly synchronized. In addition, in order to further eliminate the inconsistency of spatial scale caused by differences in sensor field of view or changes in shooting distance, all image frames are uniformly resampled to a standard scale with a spatial resolution of 1mm corresponding to 1 pixel, so that image data at different positions and at different times can be accurately aligned with each other in space. At the same time, the bit depth of each frame of the image is uniformly limited to 8 bits per pixel to simplify the subsequent processing of the image data and improve the consistency and efficiency of image processing. After the above process, a normalized 2C feature frame sequence that is completely aligned in space and time is formed. ; Indicates time.
[0031] Furthermore, in order to fully explore the dynamic change characteristics of the contact line target at different time points, the multi-dimensional recursive coupling analysis unit first performs the normalized 2C feature frame sequence output from the multimodal data acquisition unit. By length Build a stack of consecutive frames and use a step size of Slip. The principle of this operation is to embed time series information into the multi-channel image structure, so that each frame stack not only retains the spatial distribution characteristics of the current frame, but also introduces the temporal changes of the image at adjacent moments, thereby providing a continuous dynamic background for the subsequent depth model. Subsequently, to ensure that the pixel grayscale distribution within the image frame has a uniform dynamic range and sufficient contrast, a brightness linear stretch is performed on each visible light image frame, mapping the darkest pixel grayscale in the original image to 0 and the brightest pixel grayscale to 255. This linear mapping process can enhance the response of the contact line edge structure and improve the clarity of subsequent feature extraction.
[0032] For each infrared image frame, considering that its grayscale value reflects the thermal radiation intensity of the target area, it is necessary to perform a temperature range remapping operation to map the lowest radiation value in the image to 0 and the highest radiation value to 255, so that the thermal field information and the visible light image have a consistent pixel value range, which is convenient for fusion analysis. After completing the brightness and temperature normalization processing, in order to ensure the correspondence between the two modal images in spatial position, it is necessary to perform a homography transformation on the visible light image frame and the infrared image frame. Through the known camera calibration parameters and image mapping matrix, each pixel is subjected to a perspective geometric transformation to form a two-dimensional alignment grid that is aligned pixel by pixel. The core function of this alignment grid is to make any row and column index position The visible light pixels and infrared pixels on the image point to the same physical point in space, ensuring that no spatial errors are introduced when the two image modalities are fused. This processing flow enables the multi-dimensional recursively coupled parsing unit to obtain input data with uniform grayscale and precise spatial correspondence before constructing the frame stack and performing feature extraction. This provides a foundation for spatial analysis and temporal modeling of contact line states.
[0033] Furthermore, to achieve multimodal fusion enhancement of contact line image features and improve the system's recognition and positioning accuracy of contact line targets in complex environments, the multimodal data acquisition unit, after completing the homography transformation and pixel-level alignment of the visible light image frame and the infrared image frame, performs a fusion operation on each pixel position in the aligned grid to generate a fused pseudo-color tensor. The core of this fusion process lies in constructing a color representation containing three-channel information. This structurally transforms the single-channel image, which originally expressed brightness and thermal radiation characteristics separately, into a more expressive multi-channel tensor, thereby providing a more complete feature input for the deep learning model.
[0034] Specifically, for each row and column index position in the alignment grid First, the pixel brightness value of the visible light image at that position is extracted, and this value is directly used as the first color component of the fused pseudo-color tensor. This component mainly reflects the intensity change of the target under visible light conditions, retains the detailed texture, structural contours and edge features in the scene, and helps to describe the geometric shape of the contact line. Then, the grayscale value of the corresponding position is read from the infrared image. The grayscale value originally reflects the radiation intensity and is usually not in the same grayscale dynamic range as the visible light image. Therefore, it needs to be remapped to The linear interpolation operation essentially maps the original infrared grayscale values to new values proportional to their maximum and minimum values, thereby enhancing the thermal contrast of the target area in the infrared image. This remapped infrared grayscale value is then used as the second color component of the fused pseudo-color tensor. This component reflects the difference in thermal radiation between the target area and the background, making it particularly useful for contact line identification in low-light or high-contrast conditions.
[0035] To further enhance the structural contrast of the image and uniformly reflect the complementary characteristics of visible light and infrared information, the system design introduces a new third color component between the first and second color components. This third color component is generated by simple arithmetic averaging, that is, the first and second color components are numerically averaged. This operation not only plays a role in fusion in terms of numerical value, but also visually enhances the overall contrast between the contact line area and the background area, making the edge contour of the contact line more prominent, thereby providing clearer input data support for subsequent feature extraction and target positioning processes. Through this three-channel construction mechanism, the structural texture information in the visible light image and the thermal distribution information in the infrared image are complementary and fused at the pixel level, effectively improving the image's expression capabilities in terms of both spatial details and radiation characteristics.
[0036] In the organization of the fused pseudo-color tensor, the three color components of each pixel at the corresponding position are spliced at the channel level using the standard row priority order to form a fused pseudo-color tensor. The shape of this tensor is ,in Represents the number of pixel rows, Indicates the number of pixel columns, and 3 indicates the number of channels. Each tensor element It contains three sub-components, corresponding to brightness, heat and fused average value respectively. Its structural design takes into account both data compatibility and deep network input requirements, and can be directly passed as input to the subsequent multi-dimensional recursive coupling analysis unit for further multi-scale feature extraction and contact line analysis.
[0037] Furthermore, in a real-time railway 2C analysis system based on deep learning, to achieve high-precision extraction of contact line features from fused pseudo-color tensors, a multi-dimensional recursively coupled parsing unit is designed to process the input image step by step using a multi-scale convolution kernel group to extract gradient and phase features at different spatial scales. This processing strategy is based on the fact that contact lines appear as linear features with directional and elongated structures in images. Their width, edge sharpness, and background contrast vary significantly under different shooting conditions. A single-scale convolution kernel is unable to fully identify contact line structures in all their forms. Therefore, a multi-scale filtering mechanism is needed to cover both local and global image features.
[0038] Specifically, the multi-scale convolution kernel group used by the multi-dimensional recursive coupling parsing unit consists of three groups of convolution kernels of different sizes, the sizes are 、 and Each set of convolution kernels contains a transverse direction operator and a longitudinal direction operator, which can perform transverse and longitudinal differential operations on each channel in the fused pseudo-color tensor. This directional processing can enhance the response of edge structures in the image, especially for linear targets such as contact lines, whose ductility along the length direction and strong gradient characteristics in the transverse direction can be accurately expressed by directional derivatives. First, use The convolution kernel group performs preliminary local gradient extraction on the image. This scale is suitable for capturing small edge structures and helps to depict subtle brightness and heat changes around the contact line. The convolution kernel group acts on the previously obtained gradient magnitude map and phase angle map to perform a reconvolution operation on the mid-scale features, thereby enhancing the coherence of the edge area and suppressing local noise. The convolution kernel group completes the fusion of large-scale information, further smoothes the gradient and direction changes in a wide area, and highlights the structural properties of the overall coherence of the contact line in the image.
[0039] Through multi-scale progressive processing, gradient magnitude maps and phase angle maps of different levels are obtained at each scale. A cross-scale fusion operation is then performed to retain the maximum gradient magnitude across the three scales at each pixel location, while the corresponding phase angles are arithmetic averaged to construct global gradient magnitude and phase angle maps. This fusion mechanism effectively enhances the response stability of the contact line under complex backgrounds while preserving local sensitive features. Finally, the global gradient magnitude and phase angle maps are concatenated in the channel dimension according to row priority, forming a gradient phase matrix containing directional and intensity information. This provides an accurate and structurally semantically rich input foundation for subsequent morphological skeletonization and guide curve tracing.
[0040] Furthermore, in the real-time railway 2C analysis system based on deep learning, in order to accurately extract the structural contours of the contact line area and enhance the difference between it and the background, the multi-dimensional recursive coupling analysis unit introduces a multi-scale convolution kernel group, combines the transverse direction operator and the longitudinal direction operator, and performs multi-level gradient and phase analysis on the fused pseudo-color tensor to obtain a high-dimensional feature description that reflects the local edge strength and direction in the image. is generated by the multimodal data acquisition unit, and its dimension is ,in Represents the number of pixel rows, represents the number of pixel columns, and 3 represents the number of channels, which correspond to the three components of visible light brightness, infrared grayscale remapping value, and fusion average value.
[0041] First, in the multidimensional recursive coupled analytical unit, the size of The convolution kernel group acts on the tensor , using a pixel-by-pixel step sliding method, convolution operations are performed independently on each channel of the tensor to calculate the horizontal and vertical differences. Specifically, the horizontal direction operator calculates the intensity change between adjacent columns, and the vertical direction operator calculates the intensity change between adjacent rows. Assume that on a certain channel, the pixel position The horizontal difference of , the longitudinal difference is expressed as , then the local gradient amplitude under this channel can be expressed as:
[0042] ;
[0043] in, Represents the first layer of gradient magnitude map Rank In order to further obtain the directional characteristics of the local gradient, the first layer phase angle map is calculated at the same time, and its phase angle is defined as:
[0044] ;
[0045] in, is the direction angle value of the corresponding pixel in the first layer phase angle map. This angle information can represent the direction of the edge in the image and plays a key role in subsequent guiding curve tracking. The first layer gradient amplitude map and phase angle map are respectively denoted as and .
[0046] Next, the size of The convolution kernel group slides on the fused pseudo-color tensor in the same way , but the convolution operation no longer acts directly on the original image channel, but on the gradient amplitude map obtained in the first layer and phase angle diagram Through and Perform a new directional convolution operation and recalculate the horizontal difference at the mesoscale and longitudinal differential , and get the gradient magnitude map of the second layer:
[0047] ;
[0048] Correspondingly, the second-layer phase angle diagram is:
[0049] ;
[0050] The purpose of this operation is to further perform convolution smoothing on the original edge information in a larger range, improve structural coherence, reduce edge disturbances caused by local noise, and retain the direction information of the main structure of the contact line.
[0051] On this basis, the size is The convolution kernel group slides on the tensor , for the second layer gradient amplitude map and phase angle diagram Perform the same form of reconvolution operation. Repeat the directional difference calculation to obtain the gradient magnitude map of the third layer:
[0052] ;
[0053] And the corresponding third-layer phase angle diagram:
[0054] ;
[0055] After obtaining three sets of gradient magnitude maps of different scales 、 、 and phase angle diagram 、 、 Afterwards, in order to fully integrate the edge response information of each scale, the maximum value comparison operation is performed on the gradient amplitude maps of these three layers according to the pixel position, namely:
[0056] ;
[0057] This operation retains the edge information with the strongest response at any scale, enhancing the system's sensitivity to contact line structures of different widths, brightness, and thermal characteristics. Correspondingly, an arithmetic average operation is performed on the three-layer phase angle map to mitigate local fluctuations caused by extraction angles at different scales and eliminate isolated directional disturbances. The final phase angle map is:
[0058] ;
[0059] Finally, the system obtains a set of global gradient magnitude maps in the output space and global phase angle diagram , this information combination not only accurately depicts the strength of the contact line edge, but also preserves its directional distribution. In order to facilitate subsequent processing, the two are spliced in the channel dimension according to the row priority order, and the size is obtained. The gradient phase matrix , where each pixel position of Contains two components, namely:
[0060] ;
[0061] The first channel in the matrix stores edge strength information, which is used to determine the contact line area; the second channel stores edge direction information, which provides a direction reference for subsequent skeleton extraction and guide curve tracing. This is the core output of the entire multi-dimensional recursively coupled analytical unit feature extraction process. It possesses strong robustness, high expressiveness, and directional diffraction, laying the foundation for accurate contact line modeling and dynamic deviation detection. In a real-time operating environment, this mechanism ensures that the system can stably output consistent contact line structural features under various complex backgrounds and non-ideal acquisition conditions, providing effective data support for subsequent morphological analysis, curve extraction, and risk quantification.
[0062] Furthermore, in the real-time railway 2C analysis system based on deep learning, in order to extract the structural information of the contact line target from the global gradient amplitude map, the multi-dimensional recursive coupling analysis unit designed a foreground extraction and skeleton generation method based on fixed threshold judgment and morphological processing. This method is based on the gradient amplitude intensity, separates the pixels related to the contact line in the image from the background, and obtains a single-pixel width skeleton map through a step-by-step morphological processing process, providing a clear structural path for subsequent guide curve tracking and geometric deviation analysis. The core principle of this method is: the gradient amplitude map The pixel value in reflects the intensity of the grayscale change at the corresponding image position. For linear targets with clear edges and continuous structures such as contact lines, their gradient amplitude values are usually significantly higher than those of the surrounding background areas.
[0063] Therefore, first in the global gradient magnitude map The fixed threshold method is used for binarization. Set a predefined threshold , for each pixel position Perform the following judgment:
[0064] ;
[0065] in, is the initial binary image, where a pixel value of 1 indicates that the location is determined to be foreground, and a value of 0 indicates background. This initial binary image usually contains the main contact line target area, but may also contain pseudo foreground points caused by image noise or local texture.
[0066] In order to remove these non-target areas and improve the structural clarity, the multi-dimensional recursive coupling analysis unit then A morphological opening operation and a closing operation are performed sequentially. The opening operation eliminates small isolated foreground regions, while the closing operation fills small holes in the foreground, thereby improving the overall connectivity of the binary image. In particular, foreground regions smaller than 4 pixels that remain after the opening operation are further cleared to prevent these isolated regions from introducing errors in the subsequent skeleton extraction.
[0067] After the processing, the binary image After that, the system enters the morphological skeletonization stage. The purpose of this process is to compress the contact line area with a width of multiple pixels into a central path with a width of only 1 pixel to accurately reflect the direction of its geometric center. The skeletonization adopts a thinning algorithm based on iterative deletion. The core idea is to take each foreground pixel as the center and The neighborhood is judged to see whether it meets the set elimination conditions. In each iteration, the pixels that meet the conditions are marked as deletable areas, and then continue to the next round of traversal after deletion.
[0068] The specific culling rules are executed sequentially based on the structural element template, determining whether a pixel is a boundary pixel and whether its removal will not disrupt the connectivity of the current foreground area. After each round of traversal, the system checks whether there are any new removable pixels in the current binary image. If so, the next round of culling continues. If no pixel changes are produced in two consecutive rounds of iteration, it indicates that the skeletonization process has reached a stable state, that is, the binary image has converged to a single-pixel width structure.
[0069] The final skeleton diagram It is a sparsely distributed binary image, where each pixel with a value of 1 represents the skeleton center path of the contact line. This skeleton image not only has high-precision structural representation capabilities, but also has strong directional continuity and is easy to track, making it suitable for subsequent guidance curve extraction and calculation of vertical and horizontal deviations. By combining the above-mentioned fixed threshold judgment, morphological opening and closing operations, and iterative skeletonization process, the multi-dimensional recursively coupled analytical unit can achieve high-precision contact line structure recognition under unsupervised conditions, significantly improving the system's robustness and adaptability in complex environments.
[0070] Furthermore, in a real-time railway 2C analysis system based on deep learning, to accurately quantify the geometric state of the contact line, the multidimensional recursively coupled analytical unit, after completing the gradient and phase feature extraction of the fused pseudo-color tensor, threshold segmentation of the global gradient amplitude map, and morphological skeletonization, needs to further extract the center path of the contact line based on the skeleton image. Based on this center path, the vertical and horizontal deviations of the contact line relative to the design position at the current moment are calculated. This deviation calculation process not only requires the skeleton structure to have good connectivity and directional continuity, but also requires that the path extraction process accurately preserve the main structure and eliminate interference from non-target branches to improve the stability and accuracy of the deviation estimation.
[0071] First, the multi-dimensional recursive coupling parsing unit is Perform the structure scanning operation in the skeleton image for each pixel Perform degree statistics. The so-called degree refers to the pixel's The number of other skeleton pixels connected in the neighborhood. If the degree of a skeleton pixel is equal to 1, it is considered an endpoint and recorded as a set ,in For the The system then performs path tracing for each pair of endpoints, identifying the skeleton branches connected to the endpoints and calculating the pixel length of each path. The longest branch path connected to any endpoint is retained, while all other paths connected to the endpoints but less than 15 pixels in length are considered pseudo-branches and deleted. This operation effectively eliminates non-main paths caused by image noise or edge artifacts, ensuring that the central path extracted subsequently represents the physically meaningful main contact line structure.
[0072] After the pseudo-branch is cleared, the one with the smallest row coordinate is selected from the remaining endpoints as the starting point of the center path tracing, which is recorded as The physical basis of this selection strategy is that the contact line usually extends from top to bottom on the image, so the point with smaller row coordinates is more likely to be the starting direction. Based on this starting point, the system performs skeleton path tracking according to the four-connectivity rule and maintains the trajectory queue during the tracking process. , where each Represents a skeleton pixel point on the trajectory. When path tracking is initialized, the starting point Add to the track queue and mark it as visited. Then repeat the following steps until the track reaches the other endpoint: Read the current pixel from the end of the track queue ,search for skeleton pixels that have not been visited in their four-connected neighborhood. If there are multiple candidate pixels, the pixels with larger column coordinates are prioritized to ensure that the trajectory extends as far to the right as possible in the horizontal direction, thus maintaining the consistency of the path direction. Each time a new pixel is selected, it is added to the trajectory queue and its visit status is updated.
[0073] Once the trajectory queue is fully generated, the system performs a five-point sliding window center point calculation on the trajectory queue to smooth out any jitter and minor turnaround structures that may exist in the path. Specifically, the sliding window length is set to 5, and the geometric center point of every five consecutive pixels in the trajectory queue is calculated as the representative position of that interval. By sequentially connecting all center points, a preliminary smooth broken line path is formed. Because the averaging operation within the sliding window can mitigate path fluctuations caused by individual point offsets, it can improve the overall stability and consistency of the path.
[0074] For the sharp turning points in the above preliminary broken line path, angle smoothing is introduced to eliminate unreasonable path mutations. The system calculates the local angle of each section of the broken line. If the angle changes by more than 30 degrees at a certain point, it is considered an abnormal turning point. The three-point averaging method is used to correct the point, that is, the current point is replaced by the average position of the previous and next points and itself, so as to eliminate high-frequency angle mutations and improve the geometric smoothness and directional continuity of the center path. After sliding window smoothing and angle adjustment, the system finally obtains a contact line center curve with structural integrity, directional consistency and smoothness, which is recorded as the path set. , where each point Indicates the path The position of the pixel.
[0075] Finally, with the train coordinate system as a reference, the system performs vertical and horizontal projection operations on the center curve to obtain the position deviation of the contact line in the current image frame. Represents the row coordinate set of all points on the central curve, Represents a set of column coordinates. The system calculates the average of these two sets separately:
[0076] ;
[0077] in, represents the average position of the center curve in the vertical direction, Indicates its average position in the horizontal direction. Assume that the contact line design elevation preset in the system is , the design centerline position is , the vertical deviation and horizontal deviation are calculated as follows:
[0078] ;
[0079] in, Indicates vertical deviation in pixels, which can be further converted to millimeters according to spatial resolution; Represents the horizontal deviation, also in pixels. These two deviation values are the direct quantitative results of the contact line state change and will be input into the composite risk index calculation unit to participate in the calculation of the dynamic risk assessment model.
[0080] Furthermore, the composite risk index calculation unit designs a multi-factor fusion risk quantification function based on the normalized results and change trends of vertical deviation and horizontal deviation to output the composite risk index corresponding to each moment. Specifically, the composite risk index calculation unit obtains the vertical deviation corresponding to each frame of image. and horizontal deviation Then, first compare it with the system preset allowable deviation threshold and Normalization is performed. The core purpose of this normalization operation is to convert the deviation information of different units and dimensions into a dimensionless expression of a unified scale so as to perform mathematical fusion in the same risk function. Indicates time The deviation of the average position of the contact line center curve in the vertical direction relative to the design elevation, in mm; It indicates the deviation of the average horizontal position of the contact line center curve at that moment relative to the design center line, and the unit is also mm. and They are the allowable tolerances for vertical and horizontal deviations set by the system according to the contact network operation standards, and the unit is mm.
[0081] In order to comprehensively evaluate the deviation strength in the two directions, the risk function takes the fourth power of the normalized deviation ratio and then fuses them in the form of cube roots, that is, the numerator is constructed as follows:
[0082] ;
[0083] This expression structurally exhibits a nonlinear response to deviation magnitude. The fourth power amplifies the risk contribution of data with large deviations, causing the risk value to rise sharply when approaching or exceeding tolerance, thereby increasing sensitivity to potential failure conditions. The cube root operation is used to scale and balance the numerical results, ensuring that the overall function output value remains within a controllable range and avoids extreme fluctuations.
[0084] On the other hand, in order to consider the changing trend of contact line deviation in a short time, the system further introduces the instantaneous variation and , respectively represent the time With the previous moment The variation range of vertical deviation and horizontal deviation between the two directions is as follows:
[0085] ;
[0086] Normalize the instantaneous changes, take their absolute values and sum them to get the deviation change rate index:
[0087] ;
[0088] This value reflects the severity of the deviation between the current frame and the previous frame. A larger value indicates a more unstable deviation, posing a potential threat to system security. To avoid this item directly increasing the overall risk value and causing instability, the system uses the square root plus 1 to construct the denominator as an adjustment item:
[0089] ;
[0090] This structure only slightly reduces the overall risk value when the deviation changes slightly, but when the deviation fluctuates violently, it significantly increases the denominator, thereby suppressing the rising speed of the risk function and playing a dynamic regulatory role.
[0091] Finally, the overall form of the composite risk index function is defined as:
[0092] ;
[0093] in, Indicates time The composite risk index at each moment. This function has three structural characteristics: first, the numerator enhances the deviation amplitude, prioritizing responses to potential catenary structure deviations; second, the denominator regulates and suppresses short-term fluctuations to avoid system misjudgments caused by sudden fluctuations; and third, the overall function remains continuous, differentiable, and stable, making it suitable for high-frequency calculations and risk trajectory tracking in real-time systems.
[0094] This composite risk index is used in the actual system to dynamically monitor the abnormality of the contact line geometry. If the set risk threshold is exceeded, the system can automatically trigger an alarm mechanism, prompting maintenance personnel to conduct manual review or on-site inspection. Because it integrates information from two dimensions, static deviation and dynamic change, it can effectively identify the operating status of the contact line in different risk scenarios such as slight deviation, severe vibration or continuous drift, thereby providing a highly reliable and responsive intelligent analysis method for the railway power supply system, further enhancing the practical value and safety assurance capabilities of the entire deep learning-driven contact network status perception system in actual engineering applications. Low risk range: , indicating that the contact line is in a stable state, the deviation is within the design tolerance range, and the change trend is relatively gentle, so no intervention is required. Medium risk range: , indicating that the contact line deviation or change trend is close to the allowable limit, it is recommended to enter the observation state and record the data change trend. High risk area: , indicating that the contact line deviation is significant or changes dramatically, and there may be abnormal conditions such as loosening, falling off, and bracket deformation. The system should immediately alarm and prompt the operation and maintenance personnel to intervene.
[0095] The following embodiment is described using an actual train as an example. The train passes through a 10km double-track section at a constant speed of 160km / h. 100 sets of image acquisition devices arranged along the line synchronously output visible light image frames at 60fps. With infrared image frame , each frame resolution 1920×1080. Here are pixel column and row coordinates, For the The multimodal data acquisition unit is triggered by a shared 25MHz crystal oscillator, which ensures that the time drift between the visible light channel and the infrared channel is less than 0.2µs. The inertial navigation module outputs the attitude quaternion and positioning coordinates , with accuracy better than 0.02° and 0.05m respectively. Before the device goes online, the camera calibration is completed to obtain the visible light internal parameter matrix , infrared internal parameter matrix , the external parameter matrix . Calculate the homography matrix in real time during operation ,in is the rotation matrix converted from the attitude quaternion. Complete infrared alignment. The original pixel size of the two channels is 0.83mm×0.83mm, and the registration image pair is generated after resampling to 1mm:1px. To compress the illumination difference, the visible light frame performs linear stretching ,in 、 The darkest and brightest pixel grayscales of the frame; the infrared frame is normalized in the same way. Then the three channels are stitched together to generate the fused pseudo-color tensor .
[0096] at this time The three channels are the first color component, the second color component and the third color component.
[0097] Multidimensional recursive coupled analytical unit with length Window, step length Slide, build continuous frame stack Three sets of convolution kernels are applied to each tensor: sizes 3×3, 7×7, and 11×11; each set contains a horizontal operator With the longitudinal operator , Represents the kernel size. Execute the following on the three channels of the tensor: , where “*” is the convolution operation, Take the absolute value. Local gradient amplitude They are gradient amplitude map and phase angle map respectively. Take the maximum pixel value of the three-scale amplitude map to get the global gradient amplitude map , the arithmetic average of the three-scale phase images is used to obtain the global phase angle image . Splicing to generate gradient phase matrix .
[0098] exist Apply Min value :like Mark the foreground, otherwise mark the background to get the initial binary image Perform the 3×3 structural element opening and closing operation to remove noise patches smaller than 4px and produce a purified binary image. . Use Zhang-Suen iterative skeletonization: Under the odd-even two-step deletion rule, Thinning to a single-pixel skeleton When the skeleton branches are cut, scan Let the endpoints of degree 1 be the set , delete the branches connected to the endpoints and whose length is less than 15px, and keep only the trunk. Select the endpoint with the smallest row coordinate As the starting point of the trajectory, generate a trajectory sequence based on four-connected tracking . Apply 5-point window center smoothing to the series: Then, perform three-point average de-aperture on the part where the turning angle of the broken line is greater than 30° to obtain the smooth contact line center curve. .
[0099] Taking the train coordinate system as reference, The vertical projection mean coordinate is denoted as , horizontal projection average column coordinates are written as . Design elevation row coordinates , design center line coordinates Therefore, the vertical deviation .
[0100] Current frame measured 、 ,get , . Previous frame record , , so the instantaneous change The system allows the tolerance to be set to , Composite Risk Index .
[0101] Will 、 、 、 Substitute, the molecular part ; Denominator ;therefore .
[0102] in Indicates time Composite risk index; 、 is the vertical and horizontal deviation; 、 To allow for tolerance; 、 The value of 0.66 falls within the range of 0-1, indicating that the contact line is still within the tolerable range but is close to the maintenance threshold; if it continues to rise above 0.9, the system will issue an alarm and record the positioning coordinates. For maintenance decision-making. From the time the train enters the section to the current frame, the total number of processed frame pairs is GPU-measured end-to-end latency for a single frame averaged 8ms, less than the 16.7ms frame interval corresponding to 60fps, confirming that this embodiment meets real-time requirements. The RMS of deviation measurement and risk assessment errors, as measured by manual benchmarks, do not exceed 1.5mm and 0.05, respectively, demonstrating the stability and robustness of the overall process under various lighting, speed, and weather conditions.
[0103] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these specific embodiments are merely illustrative, and that those skilled in the art may omit, substitute, and modify the details of the methods and systems described above without departing from the principles and spirit of the present invention. For example, combining the above method steps to perform substantially the same functions in substantially the same manner to achieve substantially the same results falls within the scope of the present invention. Accordingly, the scope of the present invention is limited solely by the appended claims.
Claims
1. Real-time railway 2C analysis system based on deep learning, characterized by: The system includes: a multimodal data acquisition unit, a multidimensional recursive coupling analysis unit, and a composite risk index calculation unit; the multimodal data acquisition unit is used to synchronously acquire visible light image frames and infrared image frames at a fixed frame rate during train operation using image acquisition devices configured along the line, synchronously record positioning coordinates and timestamps, and perform timestamp alignment on the visible light image frames and infrared image frames to obtain a 2C feature frame sequence; the multidimensional recursive coupling analysis unit is used to construct a continuous frame stack of the 2C feature frame sequence according to a predetermined length and slide it with a set step size to regenerate a fused pseudo-color tensor; a multi-scale convolution kernel group is used to extract gradient and phase features step by step; the center curve of the contact line is determined by iterative morphological skeletonization and guide curve tracking to obtain its vertical deviation and horizontal deviation; the composite risk index calculation unit is used to calculate the composite risk index of the contact network at the corresponding time based on the vertical deviation and horizontal deviation; the multimodal data acquisition unit aligns all visible light image frames and infrared image frames with the same timestamp and uniformly resamples them to a resolution of 1mm:1px to obtain a normalized 2C feature frame sequence ; Indicates time; the bit depth of each visible light image frame or infrared image frame is 8 bits per pixel; the multi-dimensional recursive coupling analysis unit will By length Build a stack of consecutive frames and step Slide; perform linear brightness stretching on each visible light image frame, increase the grayscale of the darkest pixel to 0, and compress the grayscale of the brightest pixel to 255; perform temperature range remapping on each infrared image frame, map the lowest radiation value to 0, and map the highest radiation value to 255; perform homography transformation on the visible light image frame and the infrared image frame, align them pixel by pixel into a completely overlapping two-dimensional alignment grid, and confirm that the visible light pixel and the infrared pixel at any row and column index position point to the same spatial point; the multi-dimensional recursive coupling analysis unit performs the following operations on each pixel in the alignment grid: directly use the brightness value of the visible light pixel as the first color component; map the grayscale value of the infrared pixel to the same 0 to 255 range through linear interpolation, and then perform The second color component is constructed by taking the average value of the first color component and the second color component to enhance the overall contrast between the contact line and the background; the first color component, the second color component and the third color component are spliced at the channel level at the same row and column index position according to the row priority order to generate a fused pseudo-color tensor whose height, width and number of channels correspond to the number of pixel rows, the number of pixel columns and 3 respectively; when the multi-dimensional recursive coupling parsing unit adopts the multi-scale convolution kernel group to extract the gradient and phase features step by step, the multi-scale convolution kernel group includes three preset groups of convolution kernels with kernel sizes of 3×3, 7×7 and 11×11 respectively; each group of convolution kernels contains a horizontal direction operator and a vertical direction operator; the convolution kernel group of size 3×3 is stepped pixel by pixel The convolution kernel group of size 7×7 is slid on the fused pseudo-color tensor at a pixel-by-pixel step size, and the first-layer gradient amplitude map and the first-layer phase angle map are reconvolved. The directional difference is recalculated on the same channel to obtain the second-layer gradient amplitude map and the second-layer phase angle map. The convolution kernel group of size 11×11 is slid on the fused pseudo-color tensor at a pixel-by-pixel step size, and the second-layer gradient amplitude map and the second-layer phase angle map are convolved. The third-layer gradient amplitude map and the third-layer phase angle map are output. The corresponding relationship is established, and the maximum value of the three-layer gradient amplitude map is compared at the same position, and the maximum gradient amplitude is retained. At the same time, the arithmetic average of the three-layer phase angle map is taken to eliminate isolated directional fluctuations; the global gradient amplitude map and the global phase angle map are obtained in the same output space; according to the row priority order, the global gradient amplitude map and the global phase angle map are spliced in the channel dimension to generate a gradient phase matrix with the size of the number of pixel rows, the number of pixel columns and 2; the multi-dimensional recursive coupling analytical unit uses a fixed threshold method to determine whether the gradient amplitude value of each pixel exceeds the set threshold in the global gradient amplitude map, and marks the pixels exceeding the threshold as foreground, and the pixels below or equal to the threshold as background, thereby generating an initial binary image containing only two types of values: foreground and background;Subsequently, an opening operation and a closing operation are performed on the initial binary image in sequence to remove isolated foreground areas with an area of less than 4 pixels, and then a skeleton image is obtained through iterative morphological skeletonization, including: with the pixel as the center, the area with an area of less than 4 pixels is eliminated in the set structure element order in a three-by-three neighborhood; after completing one traversal, it is determined whether there are still deletable areas in the initial binary image; if so, the next traversal is continued; when no change occurs in two consecutive traversals, the skeletonization process converges and a skeleton image with a single pixel width is obtained; the process of obtaining its vertical deviation and horizontal deviation by the multi-dimensional recursive coupling parsing unit includes: scanning the skeleton image, identifying pixels with a degree of 1 as endpoints, retaining the longest branch associated with the endpoint, and deleting all other branches with a length of less than 15 pixels; among the skeleton endpoints, the one with the smallest row coordinate is selected as the starting point; based on the four-connected rule, its adjacent skeleton pixels are sorted in ascending order of column coordinates. Write the pixels sequentially into the trajectory queue as the initial tracking path; repeat the following operations until the other endpoint is reached: read the current pixel from the end of the trajectory queue, search for unvisited pixels adjacent to the current pixel in its neighborhood, and if multiple candidates exist, prioritize the pixel with the larger column coordinate, add the selected pixel to the trajectory queue, and mark it as visited; calculate the center point of the complete trajectory queue using a five-point sliding window, and connect the centers of consecutive windows into a smooth polyline; for locations in the polyline where the angle changes by more than 30 degrees, use the three-point averaging method to eliminate sharp turns to obtain the final contact line center curve; using the train coordinate system as a reference, project the final contact line center curve onto the vertical and horizontal axes according to the row and column coordinates: calculate the average value of the row coordinates of all pixels on the final contact line center curve, and the difference between the row coordinates and the design elevation is the vertical deviation; calculate the average value of the column coordinates of all pixels on the center curve, and the difference between the column coordinates and the design center line is the horizontal deviation.
2. The real-time railway 2C analysis system based on deep learning according to claim 1 is characterized in that: The composite risk index calculated by the composite risk index calculation unit is: ; in, For time Composite risk index at the time of For the moment Vertical deviation, unit: mm; For the moment Horizontal deviation, unit: mm; is the vertical tolerance, unit is mm; is the horizontal tolerance, in mm; is the instantaneous change of vertical deviation, unit is mm; It is the instantaneous change of horizontal deviation, in mm.
3. A real-time railway 2C analysis method based on deep learning for implementing the system according to any one of claims 1 to 2, characterized in that: The method comprises: during train operation, image acquisition devices arranged along the line synchronously acquire visible light image frames and infrared image frames at a fixed frame rate, synchronously record positioning coordinates and timestamps, and align the timestamps of the visible light image frames and the infrared image frames to obtain a 2C feature frame sequence; constructing a continuous frame stack of the 2C feature frame sequence according to a predetermined length and sliding it with a set step size to regenerate a fused pseudo-color tensor; using a multi-scale convolution kernel group to extract gradient and phase features step by step; determining the center curve of the contact line through iterative morphological skeletonization and guide curve tracking to obtain its vertical deviation and horizontal deviation; and calculating the composite risk index of the contact network at the corresponding moment based on the vertical deviation and the horizontal deviation.
Citation Information
Patent Citations
Method for measuring the uplift of electrical contact lines on rail lines
EP2942230A1
Automatic detection method of conductor height and pull-out value of overhead line system based on vehicle-mounted mobile laser point cloud
WO2023019709A1