Modular intelligent analysis method for suspension state of overhead line system
By employing a modular intelligent analysis method, the problem of unreliable analysis in complex scenarios of the overhead contact line intelligent monitoring system is solved, achieving efficient data quality control and multimodal fusion, and providing high-confidence operation and maintenance support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA RAILWAY SHENYANG BUREAU GRP CO LTD CHANGCHUN HIGH-SPEED RAILWAY INFRASTRUCTURE SECTION
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-12
AI Technical Summary
Existing intelligent overhead contact line monitoring systems provide unreliable analysis results in complex real-world scenarios, are unable to adaptively optimize, and suffer from issues such as missing data quality loops, coarse-grained state perception information, and a disconnect between multimodal fusion and other problems.
A modular intelligent analysis method is adopted, including an image pre-screening module, a key component precise positioning module, a multimodal state discrimination module, and a defect suspicion assessment and feedback module. Through image quality initial screening, component precise positioning, multimodal data fusion, and uncertainty quantification, a closed-loop optimization is formed.
It improves the reliability and adaptability of analysis results, eliminates low-quality data pollution, provides high-confidence operation and maintenance decision support, and realizes adaptive optimization of the system.
Smart Images

Figure CN122020556A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent operation and maintenance technology for rail transit, specifically to a modular intelligent analysis method for the suspension status of overhead contact lines. Background Technology
[0002] The condition of the overhead contact system in electrified railways directly determines the quality of current collection by the pantograph and the safety of train operation. Intelligent monitoring technology based on computer vision and multi-sensor fusion has become the development direction of operation and maintenance systems. However, current mainstream technical solutions have deep limitations in architectural design and physical modeling, making it difficult to meet the reliability, accuracy, and adaptability requirements of high-grade railway operation and maintenance in complex real-world scenarios. Specifically, these limitations are as follows: The lack of a closed-loop data quality management system means that existing systems generally lack effective quality assessment and screening mechanisms for raw sensor data (especially images). During inspection, motion blur, drastic changes in lighting, and temporary deviations of the target area from the field of view inevitably generate a large amount of low-quality or even invalid data. Although existing technologies (such as CN120431513A) perform out-of-focus frame removal in the later stages, they cannot prevent low-quality data from entering the core analysis process, resulting in wasted computing resources and systematically contaminating the analysis results, creating an inherent "garbage in, garbage out" problem.
[0003] State-aware information is coarse-grained. The core state of overhead contact line suspension components (such as droppers and positioners) is reflected in the deformation of their centerline geometry (such as offset, bending, and slack). Existing visual inspection methods mostly follow a general object detection framework, only outputting two-dimensional bounding boxes (such as CN120431513A) or macroscopic overall deformation functions (such as the displacement function in CN118797532A). Bounding boxes cannot characterize the continuous deformation of one-dimensional structures, and macroscopic functions cannot locate the microscopic state of specific components, resulting in a fundamental lack of information regarding the geometric features upon which subsequent judgments rely.
[0004] Multimodal fusion is disconnected from physical reality. To improve confidence levels, the introduction of sensors such as mechanical and temperature sensors for multimodal fusion has become a trend.
[0005] Existing fusion strategies are mostly shallow data stacking or feature splicing. For example, CN118797532A directly maps heterogeneous data to a health index through complex mathematical transformations, but its transformation process lacks clear physical semantics; CN120431513A simply couples visually inferred physical quantities with image features. Neither of these approaches explicitly models or compensates for known physical interferences, such as how ambient temperature significantly alters contact wire tension through thermal expansion. Without correcting for such deterministic physical relationships, directly fusing the original sensor data will inevitably lead to misjudgments of the state, and the fusion effect may even be worse than that of a single modality.
[0006] Existing solutions are all open-loop systems that output a status determination result, but cannot assess the reliability of that result itself. The system cannot distinguish between "definitely no defects" and "cannot be determined due to poor data quality," making it difficult to support high-confidence operational decisions. Furthermore, the system lacks the ability to trace the root cause of problems back from the analysis results and dynamically optimize front-end perception strategies, making it unable to adapt to constantly changing and complex environments. Summary of the Invention
[0007] The purpose of this invention is to provide a modular intelligent analysis method for the suspension state of overhead contact lines, which solves the fundamental technical problems in existing intelligent monitoring systems for overhead contact lines, such as unreliable analysis results and inability to adaptively optimize in complex real-world scenarios due to architectural coupling and lack of physical modeling.
[0008] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A modular intelligent analysis method for the suspension state of overhead contact lines includes the following steps: The image pre-screening module performs initial screening of imaging quality and determination of effective regions for the original image sequences collected by the track inspection platform, and outputs binary image validity labels. The key component precise positioning module receives images with valid labels, performs high-precision target detection to separate the core suspension components in the catenary, and outputs structured data containing component type identifiers, two-dimensional bounding box coordinates, centerline polyline descriptions, and local scale factors. The multimodal state discrimination module receives the structured data and fuses the synchronously acquired non-visual sensor data to comprehensively determine the physical state of each component, and outputs a state determination result including component identification, state category and fusion confidence value. The uncertainty of the state determination result is quantified by the defect suspicion assessment and feedback module, and a feedback signal for system closed-loop optimization is generated based on the quantification result.
[0009] In one embodiment of the present invention, the image pre-screening module calculates three quality indicators: illuminance index, high-frequency energy attenuation ratio, and spatial coverage integrity index; The illumination index is calculated by using a 5×5 pixel sliding window on the green channel of the image. When the average pixel value of the effective area of the whole image on the green channel is less than 30, it is judged as insufficient illumination. The high-frequency energy attenuation ratio is calculated by using discrete cosine transform to determine the energy percentage of DCT subbands with row and column indices both greater than 8. If this percentage is less than 15%, it is considered that the motion blur exceeds the standard. The spatial coverage integrity index is calculated by matching the feature points of the current image frame with the standard scene template to determine the intersection-union ratio and centroid offset. If the intersection-union ratio is less than 0.7 or the centroid offset exceeds 5% of the image width, it is considered that the region is missing. If any index fails to meet the preset conditions, an invalid label is output.
[0010] In one embodiment of the present invention, the key component precise positioning module adopts a two-level cascaded neural network architecture. The first level is an improved region proposal network, whose anchor frame is customized according to the aspect ratio distribution of typical catenary components and introduces an attention mechanism. The second level is a fine-tuning positioning network, whose backbone network integrates three stacked deformable convolutional layers. The offset of each level is predicted by an independent lightweight sub-network, and the input of the sub-network is the concatenation of the feature map of the previous level and the region proposal features.
[0011] In one embodiment of the present invention, the output head of the fine-tuning positioning network includes four parallel branches: The first branch outputs the probability distribution of component categories; The second branch regresses the coordinates of the four corners of the two-dimensional bounding box; The third branch generates a sequence of centerline polyline vertices consisting of 21 control points through point-by-point regression. The first and last points are constrained to the geometric centers at both ends of the component, and the middle points are obtained through interpolation. The coordinates of each point are represented in the image normalized coordinate system. The fourth branch estimates the local scale factor, which is defined as the ratio of the projected length of the target part in the image to its standard geometric model length, and scales it based on the standard length of the suspension string of 1.2 meters.
[0012] In one embodiment of the present invention, the multimodal state discrimination module synchronously acquires non-visual sensor data collected by the mechanical sensor array, temperature sensor, and environmental meteorological station deployed on the catenary support and catenary cable, and first performs temperature compensation on the mechanical sensor data. The compensation model is as follows: Corrected tension value = Measured tension value - (Current ambient temperature - Reference temperature) × Coefficient of thermal expansion × Elastic modulus × Cross-sectional area, where the coefficient of thermal expansion is 1.2 × 10⁻ 5 / ℃, elastic modulus is 180GPa, cross-sectional area is determined according to component type, reference temperature is 20℃.
[0013] In one embodiment of the present invention, in the visual domain, three geometric features, namely offset, curvature integral and endpoint spacing deviation, are calculated based on the centerline polyline description and input into a support vector machine classifier using radial basis function kernels, and the output is a visual confidence score covering four states: normal, offset, relaxed and broken. In the physical domain, the corrected tension value is compared with the lower limit of the failure threshold corresponding to the component type, and the mechanical anomaly index is calculated. When the index is less than 0.8, it is judged as abnormal. The lower limits of the tension safety thresholds for the suspension wire, positioner, and wrist arm are 8kN, 12kN, and 25kN, respectively, and each threshold is dynamically adjusted with the ambient temperature at a slope of -0.1kN / ℃.
[0014] In one embodiment of the present invention, the fusion weight of visual and physical domain evidence is determined by the type of line segment, which includes straight sections, small-radius curve sections, turnout sections, and bridge-tunnel transition sections. Each segment type corresponds to a pre-trained weight mapping function, which takes the segment's geometric parameters and historical fault statistics as inputs and outputs visual weights, with physical weights as their complements. The weight mapping function is implemented through an offline-trained multilayer perceptron with an input dimension of 7 and an output of a single visual weight value.
[0015] Furthermore, the fusion weights can be dynamically modulated based on the real-time image clarity, location reliability, and sensor signal stability to improve the robustness of the fusion decision.
[0016] In one embodiment of the present invention, the defect suspicion assessment and feedback module has a built-in Bayesian inference engine, which defines two states to be assessed: normal and defective. Its prior distribution is calculated from the frequency of occurrence of various states in historical data. Its likelihood function represents the probability of observing the current fusion confidence value in a given state, and is obtained by modeling the fusion confidence values of similar states in history through the kernel density estimation method.
[0017] In one embodiment of the present invention, the entropy value of the posterior probability distribution is calculated by Bayesian update as an uncertainty measure. When the entropy value is higher than 1.2 bits, it is determined as "insufficient data to make a judgment", otherwise it is determined as "definitely no defect" or "there is a defect".
[0018] In one embodiment of the present invention, the defect suspicion assessment and feedback module performs correlation analysis between the uncertainty measure and the quality indicators of the image pre-screening module to identify the dominant quality factors that lead to high uncertainty, and generates targeted front-end acquisition parameter adjustment instructions or sensor calibration instructions based on the analysis.
[0019] In one embodiment of the present invention, the standardized data interface adopts a binary serialization format based on Protocol Buffers, and the modules transmit data through a shared memory queue with a depth of 8 frames, and a timeout discard mechanism is configured to ensure the real-time performance of the system.
[0020] In one embodiment of the present invention, the high-speed industrial camera mounted on the track inspection platform has a frame rate of 30 frames per second, a resolution of 4096×2160, a lens focal length of 50 mm, and is equipped with a global shutter and hardware trigger synchronization mechanism; the mechanical sensor array is a fiber optic strain gauge with a sampling frequency of 1000 Hz, and the temperature sensor is a platinum resistance thermometer PT100 with an accuracy of ±0.1℃.
[0021] In one embodiment disclosed in this invention, the visual weights Before online fusion, dynamic visual weights need to be generated through dynamic modulation based on real-time credibility factors. Specifically, it includes the following steps: S131, Real-time Credibility Factor of Computer Visual Evidence : Real-time credibility factor of visual evidence The calculation is based on the global sharpness of the current image frame and the localization reliability of the currently detected part. The formula is as follows: ; This is the standard Sigmoid function; in: : Visual evidence real-time credibility factor (value range [0,1], the higher the value, the more reliable the visual evidence). The standard Sigmoid activation function is used to map input values to the [0,1] interval; : Weighting coefficient for location reliability in the visual domain, preset to a positive value, example value is 2.0; : Comprehensive positioning reliability output by the key component precise positioning module (value range [0,1]); : Weighting coefficient for visual field image sharpness, preset positive value, example value is 1.5; : The real-time sharpness score of the current image frame, with a value range of [0,1]; The input variables for the Sigmoid function; : Natural constant, approximately equal to 2.71828; Assess the real-time sharpness of the current image frame. , The high-frequency energy attenuation ratio calculated for the image pre-screening module. The function whose linear normalization is to the interval [0,1]; and These are preset positive weighting coefficients, used to adjust the location reliability. Image sharpness rating The specific values of the contribution weights to the visual credibility factor are determined through offline training or expert experience.
[0022] S132, Calculate the real-time credibility factor of physical evidence : Real-time credibility factor of physical evidence Based on the instantaneous stability of the mechanical sensor data and the reasonableness of the ambient temperature data, the calculation formula is as follows: in, For instantaneous signal-to-noise ratio estimation of mechanical sensor data, , and These are the correlation mechanical sensors in the most recent time window. Mean and standard deviation of internally sampled data To prevent division by zero of small constants; Score the reasonableness of the temperature sensor data. , This is the current temperature reading. This is the average temperature for this section of road during the same period in history. This is the preset allowable deviation threshold; and These are preset positive weighting coefficients, used to adjust the stability of the mechanical signal. With temperature rationality The specific values of the contribution weights to the physical credibility factor are determined through offline training or expert experience.
[0023] S133, Dynamic Modulation Fusion Weights: use and right Modulation is performed to obtain dynamic visual weights. : ; in: : The dynamically adjusted visual evidence weight, with a value range of [0,1]; : The basic visual weights output by the weight mapping function, with values ranging from [0,1]; To prevent extremely small constants with a denominator of zero, a minimum positive value is preset; in, To prevent extremely small constants with a denominator of zero, the dynamic physical weights are: ; Final state fusion confidence value for: ; in: The combined confidence value after fusing evidence from the visual and physical domains, ranging from [0,1], indicates that the higher the value, the greater the possibility of "the existence of a defect"; : The confidence score of the visual domain output, with a value range of [0,1], generated by the support vector machine classifier; : Confidence level of the physical domain output, with a value range of [0,1], generated by mapping the mechanical anomaly index.
[0024] In one embodiment of the present invention, the time window The time is 0.5 seconds; the weighting coefficient is: , , , .
[0025] Compared with the prior art, the present invention has the following beneficial effects: This invention establishes an independent image pre-screening module and implements hard judgments based on three quantifiable physical / geometric indicators: illuminance index, high-frequency energy attenuation ratio, and spatial coverage integrity. This creates a proactive quality barrier at the entry point of the overhead contact line visual analysis process. It eliminates invalid calculations and feature contamination caused by motion blur, insufficient lighting, and missing targets at the source, directly improving the input signal-to-noise ratio of all subsequent processing stages. This is a fundamental improvement that cannot be achieved by post-processing filtering or completely ignoring the problem in existing technologies.
[0026] This invention abandons the bounding box paradigm of general object detection and outputs a precise polyline description of the component centerline defined by 21 control points. Compared with the bounding boxes or macroscopic deformation functions provided by existing technologies, the centerline trajectory provided by this invention is the faithful geometric basis for calculating key state indicators such as offset, sag, and curvature, solving the problem of missing structural information at the source of state discrimination features.
[0027] This invention uses the coefficient of thermal expansion of materials and Hooke's law to perform temperature compensation on mechanical sensing data, eliminating deterministic interference from environmental factors and ensuring the purity of physical evidence. Through a pre-trained weight mapping function, the fusion weights of visual and physical evidence are dynamically adjusted according to the type of track section (straight line, curve, turnout, etc.). This reflects a deep understanding of the varying reliability of different sensing modalities in different environmental contexts, achieving a leap from simple data aggregation to intelligent decision-making based on physical rules and contextual knowledge.
[0028] This invention introduces an uncertainty quantification mechanism based on Bayesian inference, enabling the system to measure the confidence level of its own judgments and clearly distinguish between "certain states" and "insufficient data," providing crucial credibility for high-risk operational decisions. More importantly, it can trace high uncertainty back to specific quality indicators at the front end (such as motion blur) and generate adjustment instructions for parameters such as exposure time and acquisition frequency, forming a closed loop of "perception-analysis-optimization." This design endows the system with the ability to dynamically adapt to complex environments and continuously optimize its performance, realizing a paradigm shift from static analysis tools to adaptive intelligent systems. Attached Figure Description
[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.
[0030] Figure 1 The flowchart of the solution described in this invention is shown in the figure.
[0031] Figure 2 This is a flowchart illustrating the overall modularity of the present invention.
[0032] Figure 3 This is an architecture diagram of the precision positioning module, a key component of this invention. Detailed Implementation
[0033] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0034] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0035] Example 1: See Figures 1-3This embodiment discloses a modular intelligent analysis method for the suspension state of overhead contact lines. The method is based on an image pre-screening module, a key component precise positioning module, a multimodal state discrimination module, and a defect suspicion assessment and feedback module connected in sequence. The modules interact with each other through a standardized data interface.
[0036] The technical solution of the present invention will be described in detail below with reference to preferred embodiments.
[0037] In practical use, the image pre-screening module receives raw image sequences acquired by a high-speed industrial camera mounted on the track inspection platform. These image sequences are then coupled with spatiotemporal synchronization signals output by the inertial measurement unit and the global navigation satellite system receiver. The image pre-screening module incorporates three parallel processing subunits, used to calculate the illuminance index, high-frequency energy attenuation ratio, and spatial coverage integrity index, respectively. The illuminance index is calculated using a 5×5 pixel sliding window, performing mean filtering on the green channel of the image to obtain a local illuminance estimate. If the average pixel value of the effective area in the green channel is less than 30, it is considered insufficiently lit, triggering an invalid label.
[0038] The high-frequency energy attenuation ratio is determined by statistically analyzing the energy proportion of high-frequency coefficients after performing a discrete cosine transform on the image. The frequency band is limited to the sub-bands in which both the row and column indices of the DCT coefficients are greater than 8. When the total energy of these sub-bands accounts for less than 15% of the total energy of the entire frequency band, significant motion blur is identified, and an invalid label is triggered.
[0039] The spatial coverage integrity index is calculated by matching feature points of the current image frame with a standard scene template, calculating the intersection-union ratio (IU) of the convex hull area of the matched point set with the template region, and simultaneously detecting the centroid offset of the target region. The IU threshold is set to 0.7, and the allowable centroid offset is 5% of the image width. If the IU is below 0.7 or the centroid offset exceeds 5% of the image width, the region is considered missing, triggering an invalid label. If any of the above three indices fails to meet the preset conditions, the image pre-screening module outputs an invalid label, preventing the image from entering subsequent processing.
[0040] The critical component precise positioning module only accepts images with valid labels as input. The module employs a two-stage cascaded neural network architecture. The first stage is an improved region proposal network, whose anchor frame design is customized based on the aspect ratio distribution of typical catenary components, and an attention mechanism is introduced to enhance the response to slender structures. The second stage is a fine-tuning positioning network, whose backbone integrates deformable convolutional layers. The sampling position of the convolutional kernels is dynamically adjusted based on the coarse region output from the previous stage to adapt to non-rigid deformations of components such as droppers, positioners, and cantilever arms caused by changes in viewing angle or mechanical stress.
[0041] The deformable convolutional layers are configured as a three-level stacked structure, with the offset of each level predicted by an independent lightweight sub-network. The input to the sub-network is a concatenation of the feature map of the previous level and the region proposal features. The output head of the fine-tuned localization network contains four parallel branches: the first branch outputs the probability distribution of the part category; the second branch regresses the coordinates of the four corners of the two-dimensional bounding box; the third branch generates an ordered polyline vertex sequence of the centerline trajectory through point-by-point regression, which consists of a fixed number of control points (21 in total). The first and last points are constrained to the geometric centers at both ends of the part, and the middle points are obtained through interpolation. The coordinates of each point are represented in the image normalized coordinate system; the fourth branch estimates the local scale factor, which is defined as the ratio of the projected length of the target part in the image to its standard geometric model length. The calculation is based on the standard length of the suspension cable (1.2 meters), with other parts scaled proportionally. All outputs are integrated into a unified structured data packet by the post-processing unit. Among them, the overall localization confidence is... Defined as the product of the component category prediction probability and the detection box regression IoU (Intersection over Union) score, it is used to characterize the overall reliability of the detection. Structured data packets are passed to the next module through a standardized interface.
[0042] The multimodal state discrimination module receives structured data from the key component precision positioning module and simultaneously acquires non-visual sensor data collected by the mechanical sensor array, temperature sensor and environmental weather station deployed on the catenary support and catenary cable.
[0043] The mechanical sensor array consists of fiber optic strain gauges with a sampling frequency of 1000 Hz, and the temperature sensor is a platinum resistance thermometer (PT100) with an accuracy of ±0.1℃. The multimodal state discrimination module first performs temperature compensation on the mechanical sensor data and, based on the received line segment type data, calls a preset weight mapping function to determine the fusion weights of visual and physical evidence for the current segment. Subsequently, the multimodal state discrimination module generates independent evidence in both the visual and physical domains. The compensation model is constructed based on the material's thermal expansion coefficient and Hooke's law. Specifically, it subtracts the theoretical thermal stress (obtained by multiplying the difference between the current ambient temperature and the reference temperature by the thermal expansion coefficient, elastic modulus, and cross-sectional area of the component) from the measured tension value to obtain the corrected net mechanical response; the thermal expansion coefficient is set to 1.2 × 10⁻⁻⁻⁶. 5 / ℃, the elastic modulus is taken as 180GPa, and the cross-sectional area is determined according to the component type (for example, the cross-sectional area of the hanger can be taken as 1.0×10⁻). 6 m²), with the reference temperature set at 20℃.
[0044] Subsequently, the multimodal state discrimination module generates independent evidence in both the visual and physical domains: In the visual domain, three geometric features—offset, curvature integral, and endpoint spacing deviation—are calculated based on the centerline polyline description and input into a dedicated classifier to output visual confidence; the visual domain classifier uses a support vector machine structure with a radial basis function as the kernel function, and the training samples cover four states: normal, offset, relaxed, and fractured; In the physical domain, the corrected mechanical data is compared with the failure threshold model corresponding to the component type to calculate the mechanical anomaly index; the failure threshold model is configured differently according to the component type, with a lower limit of 8kN for the tension safety threshold of the hanger, 12kN for the positioner, and 25kN for the cantilever arm, and each threshold is dynamically adjusted with the ambient temperature at an adjustment slope of -0.1kN / ℃.
[0045] Mechanical anomaly index Defined as the corrected tension value With lower limit of safety threshold The ratio, i.e. .
[0046] The physical domain anomaly detection rules are as follows: when When this happens, it is judged as "abnormal"; when When the time is right, it is judged as "normal"; when At that time, it was determined that "further verification is required".
[0047] Physical domain confidence The mechanical anomaly index is obtained through a mapping function. The function will obtain Mapped to the [0,1] interval, the higher the value, the greater the possibility of "abnormality".
[0048] For example, a piecewise linear function can be used: ; This function is in and Continuity is ensured to guarantee a smooth transition in the physical confidence mapping.
[0049] The fusion weights of the two pieces of evidence are determined by the type of track section, which includes straight sections, small-radius curve sections, turnout areas, and bridge-tunnel transition sections. Each type corresponds to a pre-trained weight mapping function, which takes the section's geometric parameters and historical fault statistics as input and outputs weighted coefficients for visual and physical evidence. The weight mapping function is implemented by a multilayer perceptron containing two hidden layers (64 neurons per layer, ReLU activation function). Its 7-dimensional input feature vector specifically includes: track curvature radius (meters), longitudinal slope (‰), span length (meters), support type encoding (one-hot vector), frequency of faults of similar components in the section over the past year, average daily wind speed (m / s), and ambient humidity (%). The training labels are the optimal visual weight coefficients annotated through expert system retrospective analysis of historical cases. The annotation is based on the weighted ratio of visual evidence to physical evidence that maximizes the accuracy of the final state determination within that specific environment. The corresponding physical weighting coefficient is... The final state category is determined by comparing the weighted fusion confidence value with a preset threshold. The output includes the component's unique identifier, state category code, and fusion confidence value.
[0050] In the Bayesian inference engine of the defect suspicion assessment and feedback module, two states to be evaluated are defined: (Normal) and (Defect). Prior distribution Calculated from the frequency of occurrence of various states in historical data. Likelihood function. Indicates the state The current fusion confidence value was observed below. The probability is modeled as follows: collect all historically identified states. The fusion confidence value of the samples is obtained by kernel density estimation, which yields the probability density function of the confidence value in that state. ,but The posterior probability is calculated using Bayes' theorem: ; in: : The status of the overhead contact line component to be evaluated (only two types of status are defined in the document), i=0 corresponds to S0 (normal status), and i=1 corresponds to S1 (defect status). : The fusion confidence value output by the multimodal state discrimination module (value range [0,1], used to characterize the initial confidence level of the component "having a defect"); :state The prior probability is calculated from the frequency of occurrence of this type of state in historical inspection data (for example, if the normal state accounts for 95% in historical data, then P(S0) = 0.95). Likelihood function: represents the likelihood function in a given state. Under the premise of observation, the current fusion confidence value is observed. The probability is obtained by modeling the fusion confidence value of similar historical states using the kernel density estimation method; The result calculated using the total probability formula is the sum of the prior probability × likelihood function for all possible states (S0 and S1 only). This sum is used to normalize the posterior probability, ensuring the result falls within the [0,1] interval. Posterior probability, representing the probability that the component is actually in state Si after observing the current fusion confidence value c (core output result, used for subsequent uncertainty entropy value calculation).
[0051] The entropy of the posterior probability distribution is calculated as an uncertainty measure through Bayesian update, with the threshold for uncertainty entropy set at 1.2 bits. If the entropy is higher than 1.2 bits, it is determined as "insufficient data to make a judgment"; otherwise, it is determined as "definitely no defect" or "a defect exists".
[0052] Further correlation analysis was conducted between the uncertainty metric and the three quality indicators of the image pre-screening module to identify the dominant quality factors leading to high uncertainty. Targeted adjustments to front-end acquisition parameters were then generated, including exposure time correction, image acquisition frequency reconfiguration suggestions, and sensor calibration trigger signals. The exposure time correction was adjusted in steps of ±10% of the current exposure value. The image acquisition frequency reconfiguration suggestion was to select a range of 5Hz to 20Hz. The sensor calibration trigger signal included the specific sensor number and calibration mode code. These feedback signals were transmitted back to the perception and control system of the track inspection platform through a standardized control interface, forming a closed-loop optimization mechanism.
[0053] In a preferred embodiment of the present invention, the standardized data interface adopts a binary serialization format based on Protocol Buffers. Low-latency data transmission is achieved between modules through a shared memory queue with a queue depth of 8 frames. A timeout discard mechanism ensures system real-time performance. The high-speed industrial camera mounted on the track inspection platform has a frame rate of 30 frames per second, a resolution of 4096×2160, and a lens focal length of 50 mm. Combined with a global shutter and hardware-triggered synchronization mechanism, it ensures strict synchronization between image acquisition and train operation status.
[0054] In one specific embodiment, the track inspection platform operates at a speed of 80 km / h along an electrified railway trunk line. An onboard high-speed industrial camera acquires images of the overhead contact system at a rate of 30 frames per second, while an inertial measurement unit and a global navigation satellite system receiver provide precise spatiotemporal stamps. An image pre-screening module assesses the quality of each frame, eliminating images with motion blur caused by insufficient nighttime lighting, rain or fog, or missing areas due to train movement. The selected valid images are then fed into a critical component precise positioning module. This module successfully separates three core components: the dropper, the positioner, and the cantilever arm, and outputs their structured descriptive data, including category identifiers, bounding box coordinates, a 21-point centerline polyline, and local scale factors.
[0055] The multimodal state discrimination module synchronously receives data streams from fiber optic strain gauges and PT100 temperature sensors deployed on the overhead contact line supports. After temperature compensation for the tension of the dropper, it calculates a fusion confidence value by combining the geometric features of the visual domain and the mechanical response of the physical domain. The defect suspicion assessment and feedback module quantifies the uncertainty of the judgment result based on Bayesian inference. When it finds that the entropy value of a certain dropper state judgment is 1.35 bits, it determines that "data is insufficient to judge". It then analyzes the reason that the high-frequency energy attenuation ratio of the image is only 12%, triggering excessive motion blur. Therefore, it generates a feedback command to increase the exposure time by 10% and increase the image acquisition frequency to 18Hz, which is sent back to the inspection platform control system.
[0056] Comparative Example 1: The traditional single-modal visual analysis method is used, which relies solely on image data for state determination without incorporating mechanical sensing data or a temperature compensation model. When operating in the same section, this method determines the same dropper as "defective," with a fusion confidence value of 0.78, but fails to consider the thermal stress caused by the ambient temperature of 35℃ on that day.
[0057] After on-site manual verification, the dropper was found to be in normal working condition. The tension value measured by the mechanical sensor was 8.15 kN. According to the temperature compensation model of this invention, at an ambient temperature of 35°C, the corrected tension value is... Where A is the cross-sectional area of the suspension cable (in the example, 1.0 × 10⁻⁻⁴). 6 m²), calculated to Meanwhile, considering the effect of temperature on the threshold, the lower limit of the dynamic safety threshold at 35℃ is adjusted to... .
[0058] The correction value is higher than the dynamically adjusted safety threshold, meeting the safety requirements. This comparative example demonstrates that the multimodal fusion and physical model embedding mechanism proposed in this invention significantly improves the accuracy and robustness of state discrimination.
[0059] Table 1 below summarizes the performance comparison data of Example 1 and Comparative Example 1 in the same test section (50 km long, including straight sections, small-radius curve sections, and bridge transition sections): Table 1: Data shows that although the single-frame processing delay is slightly increased due to the introduction of multimodal data fusion and complex physical models, the accuracy of state discrimination is significantly improved and the false alarm rate is greatly reduced, which fully verifies the effectiveness and engineering applicability of the technical solution of the present invention.
[0060] Example 2: This example is a further optimization based on Example 1. In this example, the centerline multi-segment line description method in the key component precise positioning module makes the characterization of continuous deformation states such as suspension slack or positioner deflection more precise.
[0061] For example, for a suspension cable with an actual length of 1.25 meters, the projected length in the image is 1.18 meters, and the local scale factor is calculated to be 0.944; the polyline of its 21-point center line shows that the three middle points are obviously drooping, and the curvature integral value reaches 0.085 radians / pixel, which far exceeds the normal threshold of 0.02. Based on this, the visual domain classifier outputs a "relaxed" state with a confidence level of 0.91.
[0062] Meanwhile, the mechanical sensor measured a tension of 6.8 kN, which was corrected to 7.1 kN after compensation at an ambient temperature of 35°C, and the mechanical anomaly index was... It is in the "further verification required" range. The physical domain confidence level is calculated based on the mapping function g. The line segment is a small-radius curve segment. The weight mapping function outputs a visual weight of 0.65 and a physical weight of 0.35, with a fusion confidence value of [value missing]. If the result exceeds the 0.85 threshold, the final output is "Defect exists". The defect suspicion assessment module calculates the posterior entropy of this result to be 0.87 bits, which is lower than the 1.2-bit threshold, so it is judged as "Defect definitely exists" and a work order is generated and pushed to the operation and maintenance system.
[0063] Furthermore, the multimodal state discrimination module queries the Geographic Information System (GIS) and the track database to obtain the line section type and geometric parameters (including radius of curvature, slope, span length, support type, etc.) of the current inspection location, and combines them with environmental data such as wind speed and humidity of the day to form a 7-dimensional feature vector, which is input into a pre-trained weighted mapping multilayer perceptron to dynamically determine the fusion weights. The section geometric parameters include radius of curvature, slope, span length, support type, and historical fault frequency, totaling 7-dimensional features, which are input into the weighted mapping multilayer perceptron.
[0064] The perceptron was trained offline using 500,000 kilometers of inspection data from the past three years, employing a cross-entropy loss function and the Adam optimizer. The learning rate was set to 0.001, and the training epochs were 200, achieving a validation set accuracy of 92.3%. During online operation, the mapping function dynamically adjusts the fusion strategy of visual and physical evidence based on real-time segment information. For example, in bridge-tunnel transition sections, where vibration interference is high and mechanical signal noise is high, the physical weight is automatically reduced to below 0.3, enhancing the dominance of visual evidence. Conversely, in straight sections where mechanical signals are stable, the physical weight is increased to above 0.6, enhancing the reliability of the judgment.
[0065] The standardized data interface design ensures the decoupling and independent upgrade capabilities of each module. The message format defined by ProtocolBuffers includes image metadata, validity tags, component structured descriptions, sensor data packets, status judgment results, and feedback instructions. All fields are strongly typed, supporting version compatibility. The shared memory queue adopts a circular buffer structure. The writer appends the serialized message to the tail of the queue, and the reader consumes from the head. When the queue is full, the oldest frame is discarded, ensuring the system maintains real-time performance under high load. Real-world testing shows that on an Intel Xeon Silver 4310 processor and 64GB DDR4 memory platform, the four-stage module pipeline can stably process 30 frames per second of input, with an end-to-end latency of no more than 120 milliseconds.
[0066] This invention achieves end-to-end controllable analysis from raw sensing data to operational decision support by constructing a four-level modular processing pipeline. The image pre-screening module establishes a hard quality gate at the data entry point to prevent low-quality data from contaminating subsequent processes; the key component precise positioning module breaks through the traditional bounding box paradigm, directly extracting the geometric representation of the centerline strongly correlated with the physical state; the multimodal state discrimination module explicitly embeds a thermo-mechanical coupling physical model and dynamically adjusts the evidence fusion strategy according to the line environment; the defect suspicion assessment and feedback module endows the system with the ability to quantify its own cognitive boundaries and drives adaptive optimization of front-end sensing parameters. Each module is decoupled through standardized interfaces, ensuring both overall system synergy and support for independent iterative upgrades, providing a highly robust and accurate technical solution for intelligent operation and maintenance of overhead contact lines.
[0067] Example 3: This example aims to detail the specific implementation process, data flow, and technical effectiveness of the dynamic fusion weight adjustment method in solving the problem of instantaneous interference. This example runs on the system framework of Example 1, as follows: 1.1 Application Scenarios and Input Data: Suppose the track inspection platform is operating in a "bridge-tunnel transition section" where vibration is significant. The system is currently processing the... The frame image has been used to identify one of the target dropper components.
[0068] Input data A (from the key component precise positioning module): The overall positioning reliability of the dropper. This value is derived from the product of the softmax probability of the localization network's classification head and the IoU score of the bounding box regression.
[0069] Input data B (from the image pre-screening module): [Number] High-frequency energy attenuation ratio of frame image This value is used for motion blur determination in Example 1 (threshold 0.15); a larger value indicates a more blurred image. The normalization function... Defined as: ,in The preset fuzzy index normalization threshold is used.
[0070] Input data C (from the weighted mapping function MLP): Based on the contextual features of the "bridge-tunnel transition section" (such as radius of curvature, historical failure rate, etc.), the basic visual weights output by the MLP. Basic physical weights .
[0071] Input data D (from the visual domain classifier): Visual confidence score output by the support vector machine classifier based on the 21-point centerline geometric features of the suspension string. The indicator is "suspected slack".
[0072] Input data E (from the sensor's real-time data stream): The fiber optic strain gauge associated with this suspension cable was recently... Tension data sequence sampled within seconds: kN. Among them, "32.5kN" is an outlier value caused by instantaneous strong electromagnetic interference.
[0073] Current reading of ambient temperature sensor The database was queried to obtain the historical average temperature for this location during the same period. The system has a preset allowable deviation. .
[0074] The physical domain determination module first processes the tension data sequence using an outlier removal method based on the 3σ principle (i.e., removing data points whose difference from the mean exceeds three times the standard deviation), then evaluates the temperature-compensated and filtered data, and outputs the physical confidence score. The indicator says "Status is normal".
[0075] 1.2 Real-time credibility factor calculation process: Step 1: Calculate the real-time credibility factor of visual evidence .
[0076] First, calculate the image sharpness score. .Will Through linear functions Normalization (where 0.5 is the preset fuzzy exponential normalization threshold) yields... Therefore .
[0077] Take coefficient , Substitute into the formula to calculate: Calculations show that, although the image sharpness is generally low ( However, because the component positioning is very accurate ( The real-time credibility of visual evidence remains very high.
[0078] Step 2: Calculate the real-time credibility factor of physical evidence .
[0079] First, process the mechanical sensor data. Based on... The principle is to perform outlier removal preprocessing on the original sequence: calculate the initial mean. and standard deviation Eliminate and The difference exceeds Data points (removed in this example) Use the clean sequence after removal. Recalculate the mean and standard deviation : , .
[0080] Calculate the instantaneous signal-to-noise ratio estimate: This indicates that the current mechanical signal is completely unreliable due to interference.
[0081] Secondly, calculate the temperature rationality score. : Take coefficient , Substitute into the formula to calculate: Calculations show that due to severe interference in the mechanical signal ( Physical evidence has low real-time reliability.
[0082] 1.3 Dynamic Weight Modulation and Fusion Decision: Will , , Substitute into the modulation formula and take : ; but .
[0083] Ultimately, a fusion decision is made: ; Assume the system defect determination threshold is ,but The system ultimately determined that the dropper was in "normal condition".
[0084] 3.4 Analysis of the technical problems solved and the technical effects This embodiment reveals the decision-making risks of static fusion weights when sensors encounter strong instantaneous interference. If dynamic modulation is not used, static weights are employed. , To merge: ; Although the static fusion result (0.38) was also judged as "normal", its decision-making process unreasonably attributed physical evidence of severe interference from outliers. This results in a weighting as high as 70%. The fragility of this decision-making foundation makes it highly susceptible to misjudgment when disturbances persist or change.
[0085] The resulting technical effect: The dynamic weight adjustment mechanism automatically identifies the credibility of physical evidence. Significantly lower than visual evidence ( This dynamically reduces the weight of physical evidence from 70% to 58% and increases the weight of visual evidence from 30% to 42%. This adjustment makes the fusion decision more reliant on currently more reliable sources of evidence. The reliability of the decision outcome (normally) is enhanced because it reduces reliance on unreliable data.
[0086] Performance verification data: A comparative experiment was conducted on a test set containing 5000 frames of images (including 200 instances of sensors being subjected to transient interference). The results are shown in Table 2 below: Table 2: Experimental data confirms that the real-time reliability assessment and dynamic weight modulation method disclosed in this embodiment effectively improves the robustness and decision accuracy of the overhead contact line status analysis system under complex interference environments by following the principle of allocating decision weights based on instantaneous data quality.
[0087] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0088] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A modular intelligent analysis method for the suspension state of overhead contact lines, characterized in that, Includes the following steps: The image pre-screening module performs initial screening of imaging quality and determination of effective regions for the original image sequences collected by the track inspection platform, and outputs binary image validity labels. The key component precise positioning module receives images with valid labels, performs high-precision target detection to separate the core suspension components in the catenary, and outputs structured data containing component type identifiers, two-dimensional bounding box coordinates, centerline polyline descriptions, and local scale factors. The multimodal state discrimination module receives the structured data and fuses the synchronously acquired non-visual sensor data to comprehensively determine the physical state of each component, and outputs a state determination result including component identification, state category and fusion confidence value. The uncertainty of the state determination result is quantified by the defect suspicion assessment and feedback module, and a feedback signal for system closed-loop optimization is generated based on the quantification result.
2. The modular intelligent analysis method for the suspension state of overhead contact lines according to claim 1, characterized in that: The image pre-screening module calculates three quality indicators: illuminance index, high-frequency energy attenuation ratio, and spatial coverage integrity index. The illumination index is calculated by using a 5×5 pixel sliding window on the green channel of the image. When the average pixel value of the effective area of the whole image on the green channel is less than 30, it is judged as insufficient illumination. The high-frequency energy attenuation ratio is calculated by using discrete cosine transform to determine the energy percentage of DCT subbands with row and column indices both greater than 8. If this percentage is less than 15%, it is considered that the motion blur exceeds the standard. The spatial coverage integrity index is calculated by matching the feature points of the current image frame with the standard scene template to determine the intersection-union ratio and centroid offset. If the intersection-union ratio is less than 0.7 or the centroid offset exceeds 5% of the image width, it is considered that the region is missing. If any index fails to meet the preset conditions, an invalid label is output.
3. The modular intelligent analysis method for the suspension state of overhead contact lines according to claim 1, characterized in that: The precise positioning module for key components adopts a two-level cascaded neural network architecture. The first level is an improved region proposal network, whose anchor frame is customized according to the aspect ratio distribution of typical catenary components and introduces an attention mechanism. The second level is a fine-tuning positioning network, whose backbone network integrates three stacked deformable convolutional layers. The offset of each level is predicted by an independent lightweight sub-network, and the input of the sub-network is the concatenation of the feature map of the previous level and the region proposal features.
4. The modular intelligent analysis method for the suspension state of overhead contact lines according to claim 3, characterized in that, The output head of the fine-tuned positioning network contains four parallel branches: The first branch outputs the probability distribution of component categories; The second branch regresses the coordinates of the four corners of the two-dimensional bounding box; The third branch generates a sequence of centerline polyline vertices consisting of 21 control points through point-by-point regression. The first and last points are constrained to the geometric centers at both ends of the component, and the middle points are obtained through interpolation. The coordinates of each point are represented in the image normalized coordinate system. The fourth branch estimates the local scale factor, which is defined as the ratio of the projected length of the target part in the image to its standard geometric model length, and scales it based on the standard length of the suspension string of 1.2 meters.
5. The modular intelligent analysis method for the suspension state of overhead contact lines according to claim 1, characterized in that: The multimodal state discrimination module synchronously acquires non-visual sensor data collected by the mechanical sensor array, temperature sensor, and environmental meteorological station deployed on the catenary support and catenary cable, and first performs temperature compensation on the mechanical sensor data.
6. The modular intelligent analysis method for the suspension state of overhead contact lines according to claim 5, characterized in that: In the visual domain, three geometric features—offset, curvature integral, and endpoint spacing deviation—are calculated based on the centerline polyline description. These features are then input into a support vector machine classifier using radial basis function kernels, and the output is a visual confidence score covering four states: normal, offset, relaxed, and broken. In the physical domain, the corrected tension value is compared with the lower limit of the failure threshold corresponding to the component type, and the mechanical anomaly index is calculated. When the index is less than 0.8, it is judged as abnormal. The lower limits of the tension safety thresholds for the suspension wire, positioner, and wrist arm are 8kN, 12kN, and 25kN, respectively, and each threshold is dynamically adjusted with the ambient temperature at a slope of -0.1kN / ℃.
7. The modular intelligent analysis method for the suspension state of overhead contact lines according to claim 6, characterized in that: The fusion weights of visual and physical domain evidence are determined by the type of line segment, which includes straight sections, small-radius curves, turnout sections, and bridge-tunnel transition sections. Each segment type corresponds to a pre-trained weight mapping function, which takes the segment's geometric parameters and historical fault statistics as inputs and outputs visual weights, with physical weights as their complements. The weight mapping function is implemented using an offline-trained multilayer perceptron with an input dimension of 7 and an output of a single visual weight value.
8. The modular intelligent analysis method for the suspension state of overhead contact lines according to claim 1, characterized in that: The defect suspicion assessment and feedback module has a built-in Bayesian inference engine, which defines two states to be assessed: normal and defective. Its prior distribution is calculated from the frequency of occurrence of various states in historical data. Its likelihood function represents the probability of observing the current fusion confidence value in a given state, and is obtained by modeling the fusion confidence values of similar states in history through the kernel density estimation method.
9. A modular intelligent analysis method for the suspension state of overhead contact lines according to claim 8, characterized in that: The entropy value of the posterior probability distribution is calculated using Bayesian updates as a measure of uncertainty. When the entropy value is higher than 1.2 bits, it is determined as "insufficient data to make a judgment"; otherwise, it is determined as "definitely no defect" or "a defect exists".
10. A modular intelligent analysis method for the suspension state of overhead contact lines according to claim 9, characterized in that: The defect suspicion assessment and feedback module correlates uncertainty measurement with the quality indicators of the image pre-screening module to identify the dominant quality factors that lead to high uncertainty, and generates targeted front-end acquisition parameter adjustment instructions or sensor calibration instructions based on the analysis.