Rail transit vehicle electrical box fault prediction and health management method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO CRRC ELECTRIC EQUIP CO LTD
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-28
Smart Images

Figure CN121935684A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electrical engineering, specifically to a method and system for predicting and managing electrical box faults in rail transit vehicles. Background Technology
[0002] As the backbone of urban public transportation, the operational safety of rail transit is of paramount importance. Various electrical enclosures on the vehicles, such as traction inverter boxes and auxiliary power supply boxes, contain numerous power devices and electrical connection points, forming core components that ensure the normal operation of the vehicles. During long-term operation, these devices experience temperature rises due to factors such as current heating, increased contact resistance, component aging, and mechanical vibration. Abnormal temperature rises are an early and critical sign of electrical faults; failure to detect and address them promptly can lead to equipment damage or even safety accidents. Currently, monitoring the internal condition of rail transit vehicle electrical enclosures, especially temperature monitoring, mainly relies on traditional methods. One common method is manual periodic inspection, where maintenance personnel use handheld thermometers to measure key areas after the vehicle enters the depot. This method is not only inefficient but also heavily reliant on personnel experience, failing to capture dynamic temperature rise data during train operation and easily missing intermittent faults. Another method is to install wired temperature sensors inside the enclosures, connecting the signals to the vehicle's existing train control and management system. However, this approach involves complex wiring, is difficult to modify, and is limited by the structure and bandwidth of the onboard network. It typically only provides simple over-limit alarms, has a low data collection frequency, and lacks in-depth analysis capabilities. More importantly, existing technical solutions generally have significant drawbacks: First, their alarm thresholds are usually fixed and singular, unable to adapt to the complex and variable operating conditions of vehicles under different loads and ambient temperatures, leading to high false alarm and false negative rates. Second, their function is limited to simple judgment of the current state, belonging to passive and reactive monitoring, unable to predict equipment health trends, let alone intelligently diagnose the root causes of potential faults. As rail transit develops towards high density and long mileage, higher demands are placed on the intelligence, precision, and efficiency of operation and maintenance. Therefore, there is an urgent need in this field for an innovative technical solution that can achieve uninterrupted monitoring, possess intelligent analysis capabilities, and transform from "post-event maintenance" to "pre-event early warning" and "in-event diagnosis," fundamentally improving the safety and reliability of rail transit vehicle operation.
[0003] Therefore, the existing technology still needs further development. Summary of the Invention
[0004] The purpose of this invention is to overcome the above-mentioned technical deficiencies and provide a method and system for predicting and managing electrical box faults in rail transit vehicles, so as to solve the problems existing in the prior art.
[0005] To achieve the above-mentioned technical objectives, according to a first aspect of the present invention, the present invention provides a method for fault prediction and health management of electrical enclosures in rail transit vehicles, comprising: S1. By deploying multi-source sensors inside the electrical enclosure, at least two different types of time-series monitoring data characterizing the operating status of the enclosure are collected simultaneously; S2. Inside the electrical enclosure, using a first artificial intelligence model deployed on an edge computing unit, real-time reasoning analysis is performed on the time-series monitoring data to generate preliminary diagnostic information containing health status assessment results; S3. Based on the primary diagnostic information, dynamically adjust the abnormal judgment threshold corresponding to the current operating condition, and decide whether to trigger an early warning signal based on the comparison result between the adjusted threshold and the real-time monitoring data; S4. The primary diagnostic information from multiple electrical enclosures is transmitted to the vehicle gateway via the vehicle area network for aggregation, and a second artificial intelligence model deployed on the vehicle gateway is used for collaborative analysis to generate secondary diagnostic information for locating the specific root cause of the fault.
[0006] Specifically, the multi-source sensor includes at least a temperature sensor for monitoring the temperature of electrical connection points and a current sensor for monitoring the load current of the circuit.
[0007] Specifically, the timing monitoring data also includes vibration data collected by vibration sensors to reflect the state of mechanical components inside the enclosure.
[0008] Specifically, the first artificial intelligence model is a lightweight deep learning model that has been trained in the cloud and then distributed. The edge computing unit includes a main control microprocessor and a dedicated artificial intelligence acceleration chip that works in conjunction with the main control microprocessor.
[0009] Specifically, in step S3, the first artificial intelligence model predicts the normal fluctuation range of the enclosure temperature under the current operating conditions based on the real-time collected load current data and historical temperature data, and uses this range as the basis for the dynamic adaptive threshold.
[0010] Specifically, the health status assessment results include the predicted remaining service life of the equipment obtained based on the degradation trend analysis of the time-series monitoring data.
[0011] Specifically, in step S4, the second artificial intelligence model compares and analyzes the primary diagnostic information of devices from different enclosures but with the same function or related locations, and locates the root cause of the fault by identifying abnormal deviations in the group.
[0012] Specifically, the secondary diagnostic information is used to distinguish abnormal conditions caused by deterioration of heat dissipation, increased contact resistance at electrical connection points, or deterioration of internal component performance.
[0013] Specifically, it also includes step S5: uploading the primary diagnostic information, early warning signals and secondary diagnostic information to the ground operation and maintenance platform through the vehicle-to-ground wireless network, for generating visual reports and maintenance decision support.
[0014] According to a second aspect of the present invention, a fault prediction and health management system for electrical enclosures of rail transit vehicles is provided, comprising: An edge intelligent acquisition unit, deployed inside each electrical enclosure, includes a multi-source sensing module, an edge computing module, and a first communication module; the multi-source sensing module is used to acquire at least two different types of time-series monitoring data; the edge computing module integrates a first artificial intelligence model to process the time-series monitoring data and generate preliminary diagnostic information and early warning signals; the first communication module is used for data uploading. The vehicle-mounted intelligent gateway unit, deployed inside the vehicle, includes a second communication module and a gateway computing module; the second communication module is used to receive data from multiple edge intelligent acquisition units; the gateway computing module integrates a second artificial intelligence model for collaborative analysis of the aggregated primary diagnostic information to generate secondary diagnostic information; In addition, a vehicle-to-ground communication unit is used to realize data interaction between the vehicle-mounted intelligent gateway unit and the ground operation and maintenance center.
[0015] Beneficial effects: The fault prediction and health management method and system for electrical boxes of rail transit vehicles based on edge intelligence and multimodal data fusion provided by this invention have a series of significant positive effects compared with the prior art.
[0016] Firstly, by deploying an edge intelligent acquisition unit that integrates multiple source sensors, this invention achieves all-weather, high-precision, and multi-dimensional perception of the operating status of electrical enclosures, completely changing the outdated situation of relying on manual inspection or simple wired monitoring, and providing a rich, synchronous, and high-quality data foundation for advanced analysis.
[0017] Secondly, the core innovation of this invention lies in deploying a lightweight artificial intelligence model at the device edge, enabling real-time intelligent reasoning at the data collection source to generate the health status assessment and early warning information. This edge intelligence mode greatly reduces the dependence on cloud communication, improves the real-time performance and reliability of the system, and maintains core monitoring functions even in the event of network interruption.
[0018] Thirdly, this invention successfully solves the industry problem of inaccurate fixed threshold alarms by introducing a dynamic adaptive threshold algorithm. This algorithm can dynamically adjust the alarm threshold based on real-time load and historical data, significantly improving the accuracy of early warnings and effectively distinguishing between normal operating condition fluctuations and actual fault precursors, thereby greatly reducing false alarms and missed alarms.
[0019] Fourth, this invention transcends the scope of traditional monitoring, achieving a leap from condition monitoring to predictive maintenance. By analyzing the degradation trend of time-series data, it can make early predictions of the remaining service life of equipment, providing forward-looking quantitative basis for operation and maintenance decisions, transforming maintenance plans from passive response to proactive planning, effectively avoiding serious accidents, and improving vehicle availability.
[0020] Fifth, through the group collaborative analysis function of the vehicle gateway, this invention can compare the operating status of similar devices in the whole vehicle, accurately locate the root cause of the fault, and improve the diagnostic information from a simple "abnormality" to specific guidance on "what kind of abnormality", which greatly improves the pertinence and efficiency of maintenance work.
[0021] In summary, this invention constructs a closed-loop intelligent operation and maintenance ecosystem that integrates perception, computing, and decision-making, realizing a revolutionary transformation of the rail transit vehicle operation and maintenance mode from passive, isolated, and lagging to proactive, collaborative, and forward-looking, ultimately bringing significant value in improving safety, ensuring on-time performance, and reducing operation and maintenance costs. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the method for predicting and managing electrical box faults in rail transit vehicles provided in a specific embodiment of the present invention. Figure 2 This is a schematic diagram of the system composition of the rail transit vehicle electrical box fault prediction and health management system provided in a specific embodiment of the present invention. Detailed Implementation
[0023] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Based on the embodiments in this application, other similar embodiments obtained by those skilled in the art without creative effort should all fall within the scope of protection of this application. Furthermore, directional terms mentioned in the following embodiments, such as "up," "down," "left," and "right," are only for reference to the directions in the accompanying drawings; therefore, the directional terms used are for illustrative purposes and not for limiting the invention.
[0024] The present invention will be further described below with reference to the accompanying drawings and preferred embodiments.
[0025] Please see Figure 1This invention provides a method for predicting and managing the health of electrical enclosures in rail transit vehicles, comprising: S1. By deploying multi-source sensors inside the electrical enclosure, at least two different types of time-series monitoring data characterizing the operating status of the enclosure are collected simultaneously; It should be further explained that the core of the method of this invention lies in constructing a hierarchical and progressive intelligent diagnostic system. In specific implementation, the electrical enclosure specifically refers to the key equipment compartments such as the high-voltage box, traction inverter box, and auxiliary power supply box on rail transit vehicles. The installation of multi-source sensors must ensure the effectiveness of sensing: temperature sensors (such as PT100 platinum resistance thermometers) need to be tightly attached to the heat sink of the IGBT module with thermally conductive silicone grease; current sensors (such as closed-loop Hall effect sensors) need to be precisely fitted onto the bus of the main power circuit. The sampling frequency of all sensors is uniformly and preferably set to 10 Hz. This value is selected based on in-depth engineering analysis: the thermal inertia of rail transit electrical systems is relatively large, and its fault temperature rise process is usually on the order of seconds to minutes. A sampling rate of 10 Hz can effectively capture this dynamic process, while avoiding the massive data storage and processing burden brought by high-frequency sampling at the kilohertz level, thus achieving the optimization of system power consumption while ensuring the integrity of fault feature extraction.
[0026] Furthermore, step S1 is the foundation of system data acquisition, and its specific implementation method is as follows: 1. Sensor selection and installation specifications: ① Temperature Monitoring: The DS18B20 digital temperature sensor is preferred. This sensor is in a TO-92 package, with a measurement range of -55°C to +125°C, an accuracy of ±0.5°C within the -10°C to +85°C range, and a preferred resolution of 9 to 12 bits (corresponding to 0.5°C to 0.0625°C). Reasons for selection: Its digital output has strong anti-interference capabilities, its single-bus interface simplifies wiring, and it is cost-effective. During installation, use high-temperature resistant (>150°C), high thermal conductivity (>3W / m·K) epoxy resin adhesive (such as 3MDP420) to firmly and tightly adhere the sensor's sensing surface to the target monitoring point, such as the surface of the IGBT module's heat sink or the bolt connections of high-current busbars. To ensure temperature measurement accuracy, a thin, uniform layer of thermally conductive silicone grease (such as Shin-Etsu G-777) should be applied between the sensor's sensing surface and the metal surface to eliminate air gaps and reduce thermal resistance.
[0027] ② Current Monitoring: The ACS712ELCTR-30A-T Hall effect current sensor is preferred. This sensor is based on the Hall effect principle, has a range of ±30A, a sensitivity of 66mV / A, and its output voltage is proportional to the input current. Reason for selection: It is a non-contact measurement device, provides electrical isolation, effectively prevents common-mode interference, and has good linearity. During installation, the power line or busbar to be measured (the cross-sectional area must meet the sensor's through-hole requirements) should be passed vertically through the circular through-hole in the center of the sensor, ensuring the wire is centered within the through-hole for optimal measurement accuracy. The sensor itself should be soldered onto a small PCB board and secured to the insulating base plate of the enclosure with four M3 screws.
[0028] 2. Data synchronization acquisition mechanism: ① Hardware Triggering: To achieve strict synchronization, an advanced timer (such as TIM1) of the main control MCU (STM32H7) is used to generate a precise 10Hz pulse signal as the sampling clock. The clock source of this timer is provided by a high-precision external crystal oscillator (8MHz) after being multiplied by a PLL to ensure a stable time base.
[0029] ② Synchronization process: When an interrupt is triggered by the rising edge of the clock, the following operations are performed sequentially in the interrupt service routine: a. Start ADC conversion: Configure the MCU's internal 16-bit ADC to sample the GPIO pin connected to the ACS712 output pin. The sampled value is stored in the DMA buffer.
[0030] b. Read temperature data: Simulate single-bus timing through another GPIO pin of the MCU, send a temperature conversion command to the DS18B20, and wait for the conversion to complete (for 12-bit resolution, the maximum conversion time is 750ms, but at 10Hz sampling, there is a 100ms window, which is sufficient to complete the conversion), and then read the temperature value.
[0031] c. Timestamp: Record the precise timestamp (in Unix timestamp format, accurate to milliseconds) for the current and temperature data pairs collected at this moment, converted from the high-precision timer counter value.
[0032] ③ Data caching: The timestamped raw sensor data (current ADC value, temperature digital value) is temporarily stored in a circular buffer allocated by the internal SRAM. The buffer size can hold at least 30 minutes of data (10Hz). 60s (30min = 18000 data points) for subsequent processing.
[0033] Understandably, the above approach ensures the accuracy, synchronization, and reliability of data collection, providing a high-quality data foundation for subsequent advanced analysis.
[0034] S2. Inside the electrical enclosure, using a first artificial intelligence model deployed on an edge computing unit, real-time reasoning analysis is performed on the time-series monitoring data to generate preliminary diagnostic information containing health status assessment results; It should be further explained that the hardware core of the edge computing unit preferably consists of a heterogeneous computing architecture composed of STMicroelectronics' STM32H7A3RGT6 microcontroller (with a main frequency of up to 280MHz and a double-precision FPU) and Canaan Technology's K210 edge AI chip (with a convolutional neural network accelerator). The two communicate via a high-speed SPI interface. The first artificial intelligence model must be trained and lightweighted in the cloud before deployment.
[0035] Furthermore, step S2 is the core of edge intelligence, involving complex data preprocessing and model inference, and its specific implementation is as follows: 1. Data preprocessing pipeline: ① Current value conversion: The raw value ADC_Value (range 0-65535, corresponding to 0-3.3V) read from the ADC needs to be converted into the actual current value I (unit: A). The formula is: in, The ADC reference voltage is preferably 3.3V; The output voltage at zero current is typically Vref / 2 = 1.65V; The sensor sensitivity is 66mV / A. This formula converts the digital quantity into a physically meaningful current value.
[0036] ② Data Standardization: To meet the input requirements of the AI model, the current I and temperature T need to be standardized. The Z-score standardization method is used. in, and These are the mean and standard deviation of the training set current and temperature, calculated during the training phase. These parameters are stored as constants in Flash during model deployment.
[0037] ③ Constructing the input vector: The model input is a time series window. The preferred window length is 300 time steps (i.e., 30 seconds of historical data). During each inference, the latest 300 sets of standardized (I_norm, T_norm) data are retrieved from the circular buffer and used to construct a 300x2 matrix as the model input.
[0038] 2. The reasoning process of the first artificial intelligence model: This lightweight CNN-LSTM hybrid model has been deployed on the K210AI chip. During inference, 300x2 input data is fed into the model. The model performs the following calculations sequentially: one-dimensional convolution to extract local spatiotemporal features, ReLU activation function to introduce nonlinearity, max pooling for downsampling, LSTM layer to capture long-term dependencies, and fully connected layer for feature integration and classification. The model output is a three-dimensional vector [P_normal, P_fault1, P_fault2], representing the predicted probabilities of the three states: "normal", "increased contact resistance", and "poor heat dissipation", respectively. The sum of all probabilities is 1.
[0039] Regarding health status assessment, the category with the highest probability is taken as the preliminary diagnostic conclusion. Simultaneously, the Health Index (HI) is defined as HI = P_normal. The closer the HI value is to 1, the better the health status; the closer it is to 0, the higher the risk of failure. This HI value is the core component of the primary diagnostic information.
[0040] Understandably, the above solution provides a complete process from raw data to intelligent diagnostic results, ensuring the feasibility of edge reasoning.
[0041] S3. Based on the primary diagnostic information, dynamically adjust the abnormal judgment threshold corresponding to the current operating condition, and decide whether to trigger an early warning signal based on the comparison result between the adjusted threshold and the real-time monitoring data; It should be further explained that step S3 implements a rule-based rapid early warning mechanism that runs in parallel with the AI model, enhancing the reliability of the system. Its specific implementation method is as follows: 1. Detailed steps of the dynamic adaptive threshold algorithm: ①Historical data caching and management: A region is allocated in the external SPI Flash to store historical (current I, temperature T) data pairs for up to 7 days, along with timestamps. A first-in, first-out (FIFO) strategy is used for management.
[0042] ② Dynamic division of operating conditions: Instead of using fixed intervals, the distribution of current data in recent times (e.g., the past 24 hours) is used to automatically divide the load into three intervals (low load cluster, normal load cluster, and high load cluster) using the K-Means clustering algorithm (executed weekly at the edge). The central current value of each cluster is the representative value for that operating condition.
[0043] ③ Real-time threshold calculation: For the current value I(t) at the current moment, calculate its Euclidean distance to the center point of each cluster and assign it to the nearest cluster. Then, filter out all temperature data belonging to that cluster from historical data and calculate its mean. and standard deviation .
[0044] ④ Threshold setting and early warning: The dynamic temperature threshold is set to... The preferred value for the sensitivity coefficient k is 2.5. This choice is based on Chebyshev's inequality: for any distribution, at least (1-1 / k²) of the data falls within k standard deviations. When k=2.5, at least 84% of the data is within this range, meaning there is high sensitivity to outliers outside this range, without being overly strict to the point of generating too many false alarms, achieving the optimal engineering balance between sensitivity and specificity. If the current measured temperature T(t) > T_{threshold}, a Level 1 warning is immediately triggered.
[0045] 2. Early warning signal fusion: The system integrates the Health Index (HI) output by the AI model with the dynamic threshold warning results. The optimal warning strategy is set as follows: if HI < 0.6 or a dynamic threshold warning is triggered, a unified warning signal is generated and uploaded via the communication module. This "OR" logic ensures high reliability of the warning.
[0046] Understandably, the above scheme makes the dynamic threshold algorithm no longer a black box, and each parameter and step has a clear engineering basis.
[0047] S4. The primary diagnostic information from multiple electrical enclosures is transmitted to the vehicle gateway via the vehicle area network for aggregation, and a second artificial intelligence model deployed on the vehicle gateway is used for collaborative analysis to generate secondary diagnostic information for locating specific fault root causes. It should be further explained that step S4 realizes swarm intelligence and performs diagnosis at the system level. Its specific implementation method is as follows: 1. Data aggregation and feature construction: Each edge unit periodically (e.g., every 5 minutes) packages its primary diagnostic information (health index HI, fault probability vector, average / variance of temperature over the past 5 minutes, average current, etc.) into a data packet and sends it to the vehicle gateway via the LoRa wireless network.
[0048] After receiving data from all the same type of boxes in the vehicle (such as all 6 traction inverter boxes), the gateway first performs data alignment (based on timestamps).
[0049] Construct a comparative feature vector: For each monitored enclosure i, calculate the difference or ratio of its various indicators to the average values of other enclosures in the same group. For example: These contrastive features are combined into a new feature vector to characterize the "relative anomaly" of the box within the population.
[0050] 2. Collaborative analysis of the second artificial intelligence model: The second model is preferably an XG Boost classifier. Its input is the contrastive feature vector of each bin constructed in the previous step. After training, the model learns the unique feature patterns exhibited by different failure modes in group contrast. For example: If a box has a significantly negative Delta_HI and a significantly greater Ratio_Temp than 1, and a very high Fault1_Prob (contact resistance), the model may determine that "contact resistance has increased".
[0051] If a cabinet's Ratio_Temp is consistently high under all loads and shows a stronger correlation with ambient temperature, it may be considered to have "poor heat dissipation".
[0052] The model output is the secondary diagnostic information, which clearly points out the most likely root cause of the fault (e.g., "Traction inverter box of axle 1 of vehicle A, root cause of the fault: blockage of the heat dissipation air duct, confidence level 85%").
[0053] Understandably, the above solution provides a way to leverage differences in group data to achieve more accurate fault location and elevate diagnostic capabilities to the system level.
[0054] Specifically, the multi-source sensor includes at least a temperature sensor for monitoring the temperature of electrical connection points and a current sensor for monitoring the load current of the circuit.
[0055] It should be further noted that the DS18B20 digital temperature sensor is preferred for temperature monitoring, with a measurement range of -55°C to +125°C and an accuracy of ±0.5°C within the range of -10°C to +85°C. During installation, its metal probe should be firmly attached to the busbar connection bolts using high-temperature resistant epoxy resin to sense the true temperature of the connection point. The ACS712ELCTR-30A-T Hall effect current sensor is preferred for current monitoring, with a range of ±30A and a sensitivity of 66mV / A. During installation, ensure that the conductor under test passes through the hole in the center of the sensor and remains perpendicular to minimize measurement errors. A data synchronization mechanism is crucial: the main control MCU's timer generates a 10Hz interrupt. In the interrupt service routine, the single-bus data from the temperature sensor and the ADC value from the current sensor are read sequentially, and a uniform millisecond-level timestamp is added to this set of data. This "electro-thermal" coupled monitoring is the cornerstone of all subsequent intelligent analysis. Understandably, the beneficial effect of the above solution is that it directly obtains the original correlation data between the device's heat source (current) and its heating performance (temperature), laying the foundation for accurate modeling.
[0056] Specifically, the timing monitoring data also includes vibration data collected by vibration sensors to reflect the state of mechanical components inside the enclosure.
[0057] It should be further noted that the ADXL345 triaxial MEMS accelerometer is preferably used for vibration monitoring, with a range of ±16g and a resolution of up to 4mg / LSB. This sensor should be screwed to the inner wall of the enclosure near the cooling fan or vibration source. The vibration data sampling rate is also set to 10Hz, and communication with the MCU is achieved via an I2C bus. To extract effective features from the vibration signal, real-time signal processing is required at the edge. Specific steps include: first, removing the DC component from the raw acceleration data; then, calculating the root mean square (RMS) value of the vibration signal as a time-domain characteristic of the vibration intensity. The formula for calculating the RMS is as follows: in, This represents the root mean square value of the vibration acceleration within a single sampling window, in g. This indicates the number of sampling points within the window, preferably 100 (i.e., 10 seconds of data at a 10Hz sampling rate). This represents the acceleration value at the i-th sampling point after removing the DC component. The beneficial effect of introducing vibration monitoring is that it expands the diagnostic scope from static electrical systems to dynamic mechanical systems, effectively detecting fault modes that purely electrical sensors cannot identify, such as fan bearing wear and abnormal relay engagement, thus achieving true "mechatronics" integrated diagnostics.
[0058] Specifically, the first artificial intelligence model is a lightweight deep learning model that has been trained in the cloud and then distributed. The edge computing unit includes a main control microprocessor and a dedicated artificial intelligence acceleration chip that works in conjunction with the main control microprocessor.
[0059] It should be further explained that the construction, training, and deployment of the first artificial intelligence model are key to this invention, and the specific process is as follows: 1. Model Selection and Structure: The preferred model is a lightweight hybrid model combining a one-dimensional convolutional neural network and a long short-term memory network. Its specific structure is as follows: an input layer (receiving multimodal data from the most recent 300 time steps, i.e., a 30-second historical window); followed by two one-dimensional convolutional layers (32 and 64 filters respectively, kernel size 3, ReLU activation function); then a max-pooling layer (pooling size 2); the output is then flattened and fed into an LSTM layer (50 hidden units); finally, two fully connected layers (30 and 3 neurons respectively, output layer using Softmax activation function). This model has approximately 50,000 parameters in total.
[0060] 2. Cloud Training: Training is performed on a cloud server using a historical dataset. The dataset contains labeled data from hundreds of electrical enclosures spanning several years, with labels including "Normal," "Increased Contact Resistance," and "Poor Heat Dissipation." Before training, the data undergoes standardization preprocessing to ensure that each feature dimension has a mean of 0 and a variance of 1. Training uses the Adam optimizer with an initial learning rate of 0.001, a batch size of 64, and 100 training epochs. The training objective is to minimize the cross-entropy loss function.
[0061] 3. Model Lightweighting and Deployment: After training, the TensorFlow Lite Converter tool was used to convert the model to INT8 quantization format, compressing the model size from approximately 200KB to approximately 60KB. Finally, it was deployed to the K210AI chip at the edge via OTA (Over-The-Air).
[0062] Understandably, the beneficial effect of the above solution lies in achieving near-cloud-level complex reasoning capabilities at the resource-constrained edge through a carefully designed lightweight model and dedicated hardware, providing computing power assurance for real-time intelligent diagnosis.
[0063] Specifically, in step S3, the first artificial intelligence model predicts the normal fluctuation range of the enclosure temperature under the current operating conditions based on the real-time collected load current data and historical temperature data, and uses this range as the basis for the dynamic adaptive threshold.
[0064] It should be further explained that the specific implementation steps of the dynamic adaptive threshold algorithm are as follows. This algorithm runs independently of the AI model and provides a rule-based, interpretable, and rapid early warning layer: 1. Data Caching: A circular queue is maintained in the memory of the edge computing unit to cache the current data from the most recent 24 hours (i.e., 86,400 sampling points). and temperature Data pair.
[0065] 2. Operating condition classification: based on real-time current value The load is divided into preset load ranges. The preferred ranges are: light load (0-10A), medium load (10-20A), heavy load (20-30A), and overload (>30A).
[0066] 3. Statistical Calculation: From the cached data, identify all historical temperature data points belonging to the current load range and calculate the average value of these temperature data points. and standard deviation .
[0067] 4. Dynamic threshold setting: Calculate dynamic early warning thresholds. : in, This indicates the temperature warning threshold under the current operating conditions; This represents the average value of historical temperature data within the corresponding load range; This represents the standard deviation of historical temperature data within the corresponding load range; This is the sensitivity coefficient. The preferred value for the sensitivity coefficient k is 2.5. This value is chosen based on statistical principles: for normally distributed data, approximately 98.76% of the data points will fall within the range of... Within the range. Setting k to 2.5 means that the system will only issue warnings for extreme temperature rise events with a probability of less than 1.24%, which effectively filters out normal random fluctuations, ensures high reliability of the warnings, and achieves an optimal balance between sensitivity and false alarm rate.
[0068] 5. Warning Trigger: If the current measured temperature If this occurs, a Level 1 warning will be triggered immediately.
[0069] Understandably, the beneficial effect of the above solution is that it provides a computationally efficient, logically clear, and highly interpretable auxiliary early warning mechanism, which corroborates the diagnostic results of the AI model and together constructs a reliable early warning system.
[0070] Specifically, the health status assessment results include the predicted remaining service life of the equipment obtained based on the degradation trend analysis of the time-series monitoring data.
[0071] It should be further explained that the prediction of the remaining useful life of the equipment is achieved by a dedicated regression branch in the first artificial intelligence model. This branch uses a health index (HI) as a supervision signal during model training. The health index HI is a scalar ranging from 1 (new) to 0 (failed). For a component, its HI can be defined based on its historical performance degradation curve. For example, for an electrolytic capacitor, its HI can be correlated with the increase in its equivalent series resistance. During the inference phase, the model is input with time-series data from a recent period (e.g., the past hour) and directly outputs a current predicted health index value. RUL predictions are based on a simple linear degradation assumption extrapolated, as shown in the following formula: in, Indicates the projected remaining useful life, in hours; Defined as the failure threshold, preferably 0.3, that is, when the health index is below 0.3, the equipment is considered to have failed; This represents the current health index predicted by the model; The Health Index (HI) represents the instantaneous rate of decline in health, which can be calculated by comparing the current HI with the HI value from a certain period of time (such as 24 hours ago).
[0072] Understandably, the beneficial effect of the above solution is that it provides intuitive and quantifiable time basis for operation and maintenance decisions, realizing the leap from "status monitoring" to "life prediction".
[0073] Specifically, in step S4, the second artificial intelligence model compares and analyzes the primary diagnostic information of devices from different enclosures but with the same function or related locations, and locates the root cause of the fault by identifying abnormal deviations in the group.
[0074] It should be further explained that the second artificial intelligence model is preferably an unsupervised anomaly detection model, such as the Isolation Forest model. Specifically, the vehicle gateway collects preliminary diagnostic information from all electrical boxes of the same type (e.g., six traction inverter boxes) throughout the vehicle, including but not limited to: the probability of each category output by the model, the health index HI, and the feature layer output vector. The Isolation Forest model learns the distribution of this information under normal conditions. When the data characteristics of a certain box show a significant difference from most other boxes in the group, the model determines it as "abnormal." For example, under the same load and environment, if the temperature feature vectors of five inverter boxes cluster into one class in the feature space, while the feature vector of the sixth box is far from this cluster, the model will mark that box as abnormal and initiate root cause analysis.
[0075] Understandably, the beneficial effect of the above scheme is that it achieves fault detection "based on population normality" and also has a certain detection capability for unknown fault types.
[0076] Specifically, the secondary diagnostic information is used to distinguish abnormal conditions caused by deterioration of heat dissipation, increased contact resistance at electrical connection points, or deterioration of internal component performance.
[0077] It should be further explained that this function is achieved by training a multi-classification secondary artificial intelligence model (such as a gradient boosting decision tree model). The training features of this model include the differences, ratios, and trend characteristics of temperature, current, and vibration data between the abnormal and normal enclosures. For example: ①The characteristics of deteriorated heat dissipation conditions may be: the temperature difference between the abnormal enclosure and the ambient temperature is significantly greater than that of the normal enclosure, and its temperature is more strongly correlated with the fluctuation of the ambient temperature.
[0078] ②The characteristics of increased contact resistance may be: the absolute temperature of a specific connection point in the abnormal enclosure is much higher than that of other points, and the temperature rise has a superlinear relationship with the square of the current.
[0079] ③ The characteristics of internal component performance degradation may include: low efficiency index (if calculable) of the abnormal enclosure, or vibration signals with specific harmonic characteristics. The model output is the probability of each fault mode, such as [poor heat dissipation: 0.85, contact resistance: 0.10, component degradation: 0.05].
[0080] Understandably, the beneficial effect of the above solution is that it refines the diagnostic granularity from "there is a problem" to "what is the problem," providing direct and clear guidance for the formulation of maintenance strategies.
[0081] Specifically, it also includes step S5: uploading the primary diagnostic information, early warning signals and secondary diagnostic information to the ground operation and maintenance platform through the vehicle-to-ground wireless network, for generating visual reports and maintenance decision support.
[0082] It should be further explained that in step S5, the complete diagnostic data packet stored by the onboard gateway is uploaded to the ground maintenance platform using the WLAN or 5G network connected while the train is parked in the depot. The platform database uses a time-series database for storage. The visualization report module uses web technology (selecting ECharts) to draw a dashboard of the entire train's health status, displaying equipment health index maps, RUL prediction curves, fault alarm lists, and root cause analysis conclusions. The maintenance decision support module automatically generates an optimal maintenance plan recommendation table based on the RUL prediction results and maintenance resource availability using a queuing theory model.
[0083] Understandably, the beneficial effect of the above solution is that it realizes a complete closed loop from data collection to decision-making and application, seamlessly integrating cutting-edge edge intelligence technology into the traditional operation and maintenance management system, and maximizing its value.
[0084] Please see Figure 2 The present invention provides another embodiment, which provides a fault prediction and health management system for electrical enclosures of rail transit vehicles. The rail transit vehicle electrical enclosure fault prediction and health management system includes: An edge intelligent acquisition unit 100 is deployed inside each electrical enclosure and includes a multi-source sensing module, an edge computing module, and a first communication module. The multi-source sensing module is used to acquire at least two different types of time-series monitoring data. The edge computing module integrates a first artificial intelligence model to process the time-series monitoring data and generate preliminary diagnostic information and early warning signals. The first communication module is used for data uploading. The vehicle-mounted intelligent gateway unit 200 is deployed inside the vehicle and includes a second communication module and a gateway computing module; the second communication module is used to receive data from multiple edge intelligent acquisition units 100; the gateway computing module integrates a second artificial intelligence model for collaborative analysis of the aggregated primary diagnostic information to generate secondary diagnostic information; In addition, there is a vehicle-to-ground communication unit 300, which is used to realize data interaction between the vehicle-mounted intelligent gateway unit 200 and the ground operation and maintenance center.
[0085] Further details regarding the hardware implementation of this system are as follows: The edge intelligent acquisition unit 100 is integrated into a flame-retardant ABS plastic housing measuring approximately 100mm x 60mm x 25mm, and is mounted on the side wall of the housing via DIN rails. Its power supply is converted from the existing 24VDC power supply within the housing to 3.3V and 5V via a high-efficiency DC-DC step-down module (such as LM2596). The first communication module uses a Semtech SX1278 LoRa chip, operating at 470MHz with a transmit power of 20dBm, enabling stable communication even in the complex electromagnetic environment inside the vehicle. The onboard intelligent gateway unit 200 uses an embedded industrial computer based on an NXP MX8Mini processor, running a Linux operating system and equipped with a Python environment to run a second AI model trained using the Scikit-learn library. The vehicle-to-ground communication unit 300 directly utilizes the existing Ethernet gateway interface on the train.
[0086] Understandably, the beneficial effect of this system is that it provides a complete, reliable, and mass-producible hardware solution from sensing and computing to communication, ensuring that the aforementioned advanced methods can be executed stably and efficiently in the harsh rail transit environment.
[0087] In a preferred embodiment, this application also provides an electronic device, the electronic device comprising: The computer device includes a memory and a processor, wherein the memory stores computer-readable instructions that, when executed by the processor, implement the described method for fault prediction and health management of electrical enclosures in rail transit vehicles. The computer device can be broadly categorized as a server, terminal, or any other electronic device with the necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, memory, network interface, communication interface, etc., connected via a system bus. The processor of the computer device can be used to provide the necessary computing, processing, and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system, computer programs, etc. The internal memory can provide an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface and communication interface of the computer device can be used to connect and communicate with external devices via a network. When the computer program is executed by the processor, it performs the steps of the method of the present invention.
[0088] This invention can be implemented as a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the steps of the methods of embodiments of the invention to be performed. In one embodiment, the computer program is distributed across multiple network-coupled computer devices or processors, such that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be executed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be executed by one or more computer devices or processors, and one or more other method steps / operations may be executed by one or more other computer devices or processors. One or more computer devices or processors may execute a single method step / operation, or execute two or more method steps / operations.
[0089] Those skilled in the art will understand that the method steps of this invention can be performed by a computer program instructing related hardware, such as a computer device or processor, to perform the steps of this invention when executed. Depending on the context, any references herein to memory, storage, databases, or other media may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.
[0090] The technical features described above can be combined arbitrarily. Although not all possible combinations of these technical features are described, any combination of these technical features should be considered to be covered by this specification, provided that such combination does not contain contradictions.
[0091] The specific embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention. Any other corresponding changes and modifications made in accordance with the technical concept of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for fault prediction and health management of electrical boxes in rail transit vehicles, characterized in that, Includes the following steps: S1. By deploying multi-source sensors inside the electrical enclosure, at least two different types of time-series monitoring data characterizing the operating status of the enclosure are collected simultaneously; S2. Inside the electrical enclosure, using a first artificial intelligence model deployed on an edge computing unit, real-time reasoning analysis is performed on the time-series monitoring data to generate preliminary diagnostic information containing health status assessment results; S3. Based on the primary diagnostic information, dynamically adjust the abnormal judgment threshold corresponding to the current operating condition, and decide whether to trigger an early warning signal based on the comparison result between the adjusted threshold and the real-time monitoring data; S4. The primary diagnostic information from multiple electrical enclosures is transmitted to the vehicle gateway via the vehicle area network for aggregation, and a second artificial intelligence model deployed on the vehicle gateway is used for collaborative analysis to generate secondary diagnostic information for locating the specific root cause of the fault.
2. The method according to claim 1, characterized in that, The multi-source sensor includes at least a temperature sensor for monitoring the temperature of electrical connection points and a current sensor for monitoring the current of the circuit load.
3. The method according to claim 2, characterized in that, The timing monitoring data also includes vibration data collected by vibration sensors to reflect the state of mechanical components inside the enclosure.
4. The method according to claim 1, characterized in that, The first artificial intelligence model is a lightweight deep learning model that is trained in the cloud and then distributed. The edge computing unit includes a main control microprocessor and a dedicated artificial intelligence acceleration chip that works in conjunction with the main control microprocessor.
5. The method according to claim 1, characterized in that, In step S3, the first artificial intelligence model predicts the normal fluctuation range of the enclosure temperature under the current operating conditions based on the real-time collected load current data and historical temperature data, and uses this range as the basis for the dynamic adaptive threshold.
6. The method according to claim 1, characterized in that, The health status assessment results include the predicted remaining service life of the equipment obtained based on the degradation trend analysis of the time-series monitoring data.
7. The method according to claim 1, characterized in that, In step S4, the second artificial intelligence model compares and analyzes the primary diagnostic information of devices from different enclosures but with the same function or related locations, and locates the root cause of the fault by identifying abnormal deviations in the group.
8. The method according to claim 1, characterized in that, The secondary diagnostic information is specifically used to distinguish abnormal conditions caused by deterioration of heat dissipation conditions, increased contact resistance at electrical connection points, or deterioration of internal component performance.
9. The method according to claim 1, characterized in that, It also includes step S5: uploading the primary diagnostic information, early warning signals and secondary diagnostic information to the ground operation and maintenance platform through the vehicle-to-ground wireless network, for generating visual reports and maintenance decision support.
10. A fault prediction and health management system for electrical enclosures of rail transit vehicles, used to implement the method as described in any one of claims 1-9, characterized in that, include: The edge intelligent acquisition unit is deployed inside each electrical enclosure and includes a multi-source sensing module, an edge computing module, and a first communication module; The multi-source sensing module is used to collect at least two different types of time-series monitoring data; the edge computing module integrates a first artificial intelligence model, which is used to process the time-series monitoring data and generate preliminary diagnostic information and early warning signals; the first communication module is used for data uploading; An in-vehicle intelligent gateway unit, deployed inside a vehicle, includes a second communication module and a gateway computing module; the second communication module is used to receive data from multiple edge intelligent acquisition units; The gateway computing module integrates a second artificial intelligence model, which is used to collaboratively analyze the aggregated primary diagnostic information and generate secondary diagnostic information. In addition, a vehicle-to-ground communication unit is used to realize data interaction between the vehicle-mounted intelligent gateway unit and the ground operation and maintenance center.