Fault prediction and health management method for rail transit intelligent traction converter
By collecting and analyzing the electrical and state variables of rail transit traction converters in real time, and establishing a degraded state space model using intelligent algorithms, the problems of low fault prediction accuracy and poor real-time performance in existing technologies are solved. This enables accurate fault prediction and health management, reduces equipment maintenance costs and safety risks, and improves the system's intelligence level.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies struggle to achieve real-time fault prediction and health management of rail transit traction converters under complex operating conditions, resulting in low fault prediction accuracy and poor real-time performance. This makes it impossible to detect potential faults in a timely manner, increasing equipment maintenance costs and safety risks.
By collecting electrical and state variables of the traction converter in real time, combining data analysis with intelligent algorithms, a degraded state-space model is established, mean field variational inference and Dirichlet temperature scaling calibration are implemented, a safety barrier function is constructed, the failure probability vector and remaining lifetime are estimated, and health management instructions are optimized.
It improves the accuracy of fault prediction, enables precise assessment of the remaining service life of traction converters, adapts to complex operating conditions, reduces equipment maintenance costs and downtime, and enhances system safety and intelligence.
Smart Images

Figure CN121784418A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rail transit technology, and in particular to a method for fault prediction and health management of intelligent traction converters for rail transit. Background Technology
[0002] Traction converters are key components in modern rail transit systems, responsible for converting electrical energy from the power grid into the power required to drive the motors. Due to their complex operating environment, large load fluctuations, and long service life, traction converters are frequently affected by various factors during operation, such as current and voltage fluctuations, temperature changes, and equipment aging. These factors can lead to a gradual decline in the performance of the traction converter, and even malfunctions. Traction converter failures not only directly affect the normal operation of trains, but in severe cases can also lead to prolonged shutdowns, causing significant economic losses and safety hazards.
[0003] Traditional traction converter fault diagnosis methods typically rely on periodic manual inspections, simple electrical signal monitoring, or early warning systems based on static models. These methods suffer from slow response times, inaccurate diagnosis, and an inability to predict faults in advance, making it difficult to provide high-precision fault prediction and real-time health management in dynamic operating environments. Especially under complex operating conditions, traditional methods often struggle to cope with the dynamic operating status of traction converters, failing to accurately capture potential fault risks and resulting in a lack of timely and effective preventative measures before faults occur.
[0004] Therefore, how to achieve real-time fault prediction and health management of rail transit traction converters, especially how to accurately predict the fault occurrence time and remaining service life of traction converters based on real-time monitoring data under variable operating conditions, has become a key technical problem that urgently needs to be solved. Existing solutions in this regard generally suffer from poor real-time performance, low prediction accuracy, and inability to adapt to complex operating conditions, necessitating a more efficient and intelligent fault prediction and health management method. Summary of the Invention
[0005] To overcome the above-mentioned technical defects, the present invention aims to provide a fault prediction and health management method for intelligent traction converters in rail transit. By collecting electrical and status parameters of the traction converter in real time and combining them with intelligent algorithms for data analysis, the method can accurately predict the fault occurrence time and remaining service life, thereby achieving accurate fault prediction and health management.
[0006] This invention discloses a fault prediction and health management method for intelligent traction converters in rail transit, characterized by comprising the following steps: Step S1: Collect electrical and status variables of the traction converter control system, including DC bus voltage. Three-phase current Transient collector-emitter voltage waveform of power semiconductor switch Gate current waveform On-state voltage drop sampling value Gate control signal Device junction temperature estimate and operating condition parameters ,in For train speed, Traction / braking mode; Step S2: Under traction error constraints Below, low-amplitude diagnostic perturbations are injected into the carrier duty cycle. perturbation frequency satisfy carrier frequency Simultaneously record the voltage change caused by the perturbation. With change in current ; Step S3: In the preset frequency band Internal calculation of equivalent series resistance of DC link And extract the gate charge integral. Duration of shutting down the Miller platform and spectral sideband energy ; Step S4: Generate a directed graph of three phases-bridge arm-submodule based on the topology configuration file, construct the channel-submodule topology association matrix M, and then... Mapped to submodule-level feature vectors via M ; Step S5: Establish a degenerate state-space model, which takes the following form: ,
[0007] in, State vector; input vector And set the variable to indicate the variable. Threshold for distinguishing change points When the change point is determined by the log-likelihood ratio statistic The degradation mechanism of the degenerate statespace model changes at that time. Step S6: Perform mean-field variational inference on the degraded state-space model to obtain the failure probability vector. Remaining life estimate and remaining lifetime uncertainty ; Step S7: Perform Dirichlet temperature scaling calibration on the fault occurrence probability vector, with temperature scaling parameters... The calibrated failure probability vector is obtained. An isotropic univariate order-preserving regression calibration is performed on the remaining lifetime estimate to establish a monotonic calibration mapping. The calibrated remaining lifetime estimate is obtained. ; Step S8: Construct the safety barrier function : , , ; One of them, Within the time window The mean, Within the time window standard deviation This refers to the upper limit of the rated peak allowable current of the device. Set for peak current. For safety reasons, For the standby threshold; within the time budget Internally solve the following formula to determine the health management instruction variables. : , ; in , , For weight parameters, The modulation ratio; Step S9: Issue the health management command to the traction converter control unit and schedule the time. Internal execution; upper limit of junction temperature Given by the device's rated thermal parameters; Step S10: When the population stability index PSI satisfies At that time, incremental learning is performed and a fixed sample replay ratio is used. The sample replay strategy updates the model parameters; PSI is calculated using the binning histogram of the reference window and the current window distribution.
[0008] Preferably, the diagnostic perturbation uses the jitter mode of a pseudo-random binary sequence (PRBS). upper limit of amplitude satisfy and The duration is one carrier cycle and and satisfy .
[0009] Preferably, the spectral sideband energy In the set of sideband frequencies Above calculation; Duration of Miller platform shutdown Transient collector-emitter voltage waveform of power semiconductor switch The first-order difference zero-crossing count of the platform segment is obtained; gate charge integral. The gate current waveform during one switching cycle Points are earned.
[0010] Preferably, the device junction temperature estimate is obtained by a dual-mode thermal observer; the dual-mode thermal observer includes an equivalent Type of thermal network parameters , , , The recursive observations, and in Enable with forgetting factor Recursive least squares identification for online updating of equivalent thermal parameters , The root mean square of the observed residuals was used as the mode switching criterion.
[0011] Preferably, the channel-submodule topological association matrix M is automatically generated based on the directed graph of the phase-arm-submodule; a submodule health status vector is defined. Using group sparsity and The joint regularization degradation coefficient estimation model has the following objective function:
[0012] Wherein, the design matrix X is composed of sub-module level eigenvectors Composition, degenerate observation vector Regularization coefficient , Parameter group .
[0013] Preferably, the mean field variational inference is used for... and The joint posterior distribution is used to solve the coordinate ascent problem, and the failure probability vector is output. Remaining life estimate and its variance The log-probability vector is softmax is the log-log odds vector. The class-wise exponential normalization mapping.
[0014] Preferably, the Dirichlet temperature scaling calibration is applied to the fault occurrence probability vector. Temperature scaling parameters Anisotropic univariate ordinal-preserving regression calibration is applied to the remaining lifetime estimate. Monotonic calibration mapping is .
[0015] Preferably, the optimization of health management instructions adopts secondary planning and time budgeting. Internal solution, constraint set includes the upper limit of the device's rated peak allowable current. ,junction temperature upper limit and based on calibrated remaining lifetime estimates Disposal threshold Output health management instruction variables .
[0016] Preferably, time synchronization employs a boundary clock method based on a precision time protocol for unified channel calibration, and linear interpolation compensation is applied to the sampling timestamp to ensure that the deviation between the acquisition clock and the control clock meets the requirements. When referenced again, This is the synchronization accuracy threshold.
[0017] Preferably, when the PSI exceeds the trigger threshold And the change point discriminant log-likelihood ratio statistic Exceeding the threshold When the H-infinity robust state estimator is enabled, the robust term coefficient is... Its optimized formula is Output the probability vector of backup failure occurrence. Compared with the estimated remaining reserve life The primary and backup inferences are selected based on an uncertainty threshold; where the uncertainty radius is... .
[0018] Compared with existing technologies, the above technical solution has the following advantages: 1. Improving the accuracy of traction converter fault prediction. In existing technologies, traditional fault prediction methods typically rely on periodic inspections or simple electrical signal monitoring, failing to provide real-time, accurate fault prediction. These methods suffer from slow response times and the inability to provide effective early warnings before faults occur, leading to untimely detection of potential faults and posing significant challenges to equipment maintenance. This invention, however, collects electrical and status parameters of the traction converter in real time and utilizes intelligent algorithms for data analysis. By monitoring real-time signals and combining this with intelligent prediction models, the timing and type of traction converter faults can be accurately predicted. This technical solution effectively improves the accuracy of traction converter fault prediction, enabling timely fault prediction before they occur and allowing for preventative measures to be taken in advance, thereby avoiding sudden equipment failures and reducing downtime and maintenance costs.
[0019] 2. Achieving accurate assessment of the remaining service life of traction converters. Existing technologies typically assess the remaining service life of traction converters through simple equipment inspections or estimations based on historical data, lacking real-time analysis of equipment operating status. Therefore, it is difficult to accurately determine when the equipment will reach a critical failure state. This invention, by collecting electrical and status data of the traction converter and combining a degraded state-space model with mean-field variational inference technology, accurately estimates the remaining service life of the traction converter. This scheme performs dynamic assessment based on real-time monitoring data, taking into account the actual working environment and status changes of the equipment. This approach enables accurate assessment of the remaining service life of the traction converter, providing a scientific basis for maintenance decisions. The system can adjust maintenance plans based on real-time data, achieving more efficient preventative maintenance, extending equipment life, and reducing the occurrence of sudden failures.
[0020] 3. Adapting to complex operating conditions and improving the real-time performance and accuracy of fault prediction. Traditional fault diagnosis methods are mostly based on static models or historical data. These methods cannot adapt to the complex changes in traction converters under different operating conditions, resulting in poor accuracy and real-time performance of fault prediction under complex conditions such as load changes and temperature fluctuations. This invention introduces incremental learning and sample replay strategies, combined with real-time monitoring data and adaptive analysis of complex operating conditions, to dynamically adjust according to different operating conditions, thereby improving the real-time performance and accuracy of fault prediction. This scheme enables traction converters to maintain high-precision fault prediction capabilities under complex operating conditions, improving the adaptability and real-time performance of fault diagnosis, providing accurate early warnings under various operating conditions, and reducing the potential risk of equipment failure.
[0021] 4. Reduce maintenance costs and downtime of traction converters. Traditional fault diagnosis methods often fail to accurately predict the timing of faults, leading to frequent emergency repairs after a fault occurs, resulting in long downtime and high maintenance costs. This invention accurately predicts the fault occurrence time and remaining service life of the traction converter, allowing for advance maintenance planning and dynamic adjustments based on the actual health status of the equipment. This avoids faults or enables pre-fault maintenance preparation. The technical solution of this invention significantly reduces equipment downtime and emergency repair costs caused by sudden faults, optimizes the traction converter maintenance process, reduces equipment downtime, and improves system operational efficiency.
[0022] 5. Improve system safety and reduce accident risks. Existing technologies lack real-time health management methods and often can only take emergency measures after a failure occurs. Especially under conditions of large load fluctuations or frequent environmental changes, traditional technologies cannot effectively deal with potential equipment failures, increasing the risk of accidents. This invention, by monitoring the operating status of the traction converter in real time and combining it with an intelligent health management model, can detect potential failure risks in real time and make timely adjustments based on data to ensure that the equipment operates in a safe state. This technical solution effectively improves the operational safety of the traction converter, can promptly alarm and take corresponding measures when equipment failure signs appear, reduces the risk of accidents caused by equipment failures, and provides strong protection for the safety of rail transit systems.
[0023] 6. Enhance the system's intelligence level and improve its autonomous operation capability. Traditional fault diagnosis methods mostly rely on manual intervention, lacking the ability to autonomously judge and automatically adjust. This makes it difficult for the system to make autonomous optimization adjustments when complex operating conditions or emergency faults occur. This invention introduces intelligent algorithms and incremental learning technology, combined with real-time operating data for automatic analysis and adjustment. It can dynamically optimize operating parameters based on the system's health status without manual intervention, thus improving the system's autonomous operation capability. This technical solution gives the system a higher level of intelligence, enabling it to autonomously adjust and optimize in complex operating environments, reducing the need for manual intervention and improving the system's automation level and reliability. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating a fault prediction and health management method for an intelligent traction converter in rail transit. Detailed Implementation
[0025] The advantages of the present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments.
[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0027] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0028] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0029] In the description of this invention, it should be understood that the terms "longitudinal", "lateral", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0030] In the description of this invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection" and "linking" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two components. They can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.
[0031] In the following description, suffixes such as "module," "part," or "unit" used to denote elements are used only for the convenience of the description of the invention and have no specific meaning in themselves. Therefore, "module" and "part" can be used interchangeably.
[0032] See Figure 1 As shown, this embodiment provides a fault prediction and health management method for an intelligent traction converter in rail transit, including the following steps: Step S1: Collect electrical and status quantities of the traction converter control system, including DC bus voltage. Three-phase current Transient collector-emitter voltage waveform of power semiconductor switch Gate current waveform On-state voltage drop sampling value Gate control signal Device junction temperature estimate and operating condition parameters ,in For train speed, For traction / braking mode; Step S2: Under traction force error constraint Below, low-amplitude diagnostic perturbations are injected into the carrier duty cycle. perturbation frequency satisfy carrier frequency Simultaneously record the voltage change caused by the perturbation. With change in current Step S3: In the preset frequency band Internal calculation of equivalent series resistance of DC link And extract the gate charge integral. Duration of shutting down the Miller platform and spectral sideband energy Step S4: Generate a directed graph of three phases-bridge arm-submodule based on the topology configuration file, construct the channel-submodule topology association matrix M, and then... The M is mapped to a submodule-level feature vector. Step S5: Establish a degenerate state-space model, which takes the following form: , ,in, State vector; input vector And set the variable to indicate the variable. Threshold for distinguishing change points When the change point is determined by the log-likelihood ratio statistic Step S6: Determine that the degradation mechanism of the degraded state space model has changed; Step S7: Perform mean field variational inference on the degraded state space model to obtain the failure probability vector. Remaining life estimate and remaining lifetime uncertainty Step S7: Perform Dirichlet temperature scaling calibration on the fault occurrence probability vector, with temperature scaling parameters... The calibrated failure probability vector is obtained. Perform an isotropic univariate order-preserving regression calibration on the remaining lifetime estimate to establish a monotonic calibration mapping. The calibrated remaining lifetime estimate is obtained. Step S8: Construct the safety barrier function : , , ; one of them, Within the time window The mean, Within the time window standard deviation This refers to the upper limit of the rated peak allowable current of the device. Set for peak current. For safety reasons, For the standby threshold; within the time budget Internally solve the following formula to determine the health management instruction variables. : , ;in , , For weight parameters, For modulation ratio; Step S9: Send the health management command to the traction converter control unit and in the time budget Internal execution; upper limit of junction temperature Given by the device's rated thermal parameters. Step S10: When the population stability index PSI satisfies... At that time, incremental learning is performed and a fixed sample replay ratio is used. The sample replay strategy updates the model parameters; the PSI is calculated using the binning histogram of the reference window and the current window distribution.
[0033] In this embodiment, step S1 will be described in detail. This step involves acquiring multi-source electrical signals and status variables of the traction converter control system, specifically including the DC bus voltage. Three-phase current Transient collector-emitter voltage waveform of power semiconductor switch Gate current waveform On-state voltage drop sampling value Gate control signal Device junction temperature estimate and operating condition parameters It should be noted that this embodiment will describe five aspects: measurement chain design, synchronization and time reference, anti-aliasing and quantization accuracy, drift and self-calibration, data quality control and metadata management.
[0034] Firstly, regarding the measurement chain design, the DC bus voltage... Sampling is performed using a combined high-voltage divider and isolation operational amplifier sensor chain, with the equivalent input noise density not exceeding [a certain value]. The bandwidth covers 0 to the carrier frequency. 2.5 times, to ensure subsequent operation within the preset frequency band. The energy spectrum calculation within the system is unaffected by the front-end roll-off. Three-phase current. A dual-channel redundant structure employing a closed-loop Hall current sensor in parallel with a milliohm-level shunt resistor is used. The closed-loop Hall channel provides wide bandwidth and isolation, while the shunt channel provides high linearity and high-precision reference. The two channels are registered using least-squares within the calibration range to form a fused current estimate. This fused current estimate exhibits repeatability better than ±0.15%FS at 25℃, and a total error not exceeding ±0.5%FS within the temperature range of -25℃ to 85℃. The transient collector-emitter voltage waveform during switching is also described. Acquired using a high-speed differential probe and a 12-bit high-speed sampling channel, with a small-signal bandwidth of no less than 50MHz and an equivalent sampling interval of no more than 20ns on the rising edge, to ensure the duration of the subsequent turn-off Miller plateau. The first-order differential zero-crossing count has sufficient time resolution. Gate current waveform. The goal is to accurately calculate the gate charge integral within a single switching cycle by employing a low-resistance gate series sampling resistor and a broadband differential amplifier sampling path. The integration time window and the gate control signal The on / off timing is consistent. On-state voltage drop sampling value. The acquisition is performed in the conduction steady-state plateau region using a timed sampling + noise-resistant averaging method, with the timing controlled by the gate control signal. The rising edge is locked to avoid transient "tail current" disturbances. Device junction temperature estimate. The data is provided by a dual-mode thermal observer, where the equivalent π-type thermal network is recursively observed under normal conditions, and the log-likelihood ratio statistic is used to determine the point of change. Exceeding the threshold Switching to a model with a forgetting factor Recursive least squares identification for online updating of equivalent thermal parameters , This provides a continuous and consistent temperature measurement for constraints related to device junction temperature in subsequent safety barrier functions. Operating parameters , Speed is provided synchronously by the train network and traction control unit. Redundant source consistency verification based on speed measuring shaft or encoder is adopted for traction / braking mode. Output from the upper-level traction dynamics state machine.
[0035] Secondly, regarding synchronization and time reference, to ensure complete alignment of multi-source measurements in the time domain, this step, along with subsequent steps, uses the boundary clock scheme of the Precision Time Protocol (PTP) to establish a full-link clock domain; the timestamps of each sampling channel are aligned with a unified clock and then linearly interpolated at the data receiver to compensate for the deviation between the acquisition clock and the control clock. No more than In some embodiments, Set to 0.5 μs, after a 24-hour drift test, the 99.7 percentile... The latency remains less than 0.8 μs, which ensures the DC bus voltage... Three-phase current Transient collector-emitter voltage waveform during switching Gate current waveform With gate control signal The phase relationship can be used for subsequent assemblies based on sideband frequencies. Extracting spectral sideband energy And calculate the duration of the shut-off Miller platform. .
[0036] Furthermore, regarding anti-aliasing and quantization accuracy, this step focuses on the DC bus voltage. With three-phase current The sampling adopts a synchronous equal-interval sampling rate. A third-order Butterworth anti-aliasing filter was installed at the analog front end, with a -3 dB cutoff frequency set to 0.45 dB. In this embodiment, Taking 200 kS / s, thus for the carrier frequency Systems within the 5-15kHz range can ensure the preset frequency band. It lies entirely within the Nyquist bandwidth. To simultaneously meet the requirements of high-speed transient acquisition, the switching transient collector-emitter voltage waveform... Gate current waveform Employs an independent high-speed sampling channel with a sampling rate of A typical value is 550 MS / s, and an on-chip digital down-conversion window is used to align the gate control signal within a local time window. The information along the line is compressed to reduce the data volume. Regarding analog-to-digital conversion accuracy, the DC bus voltage... With three-phase current The quantization bit width is no less than 16 bits, and the high-speed channel bit width is no less than 12 bits; after unified gain and zero-point calibration, the DC bus voltage The three-sigma quantization noise corresponds to a voltage of no more than 0.05%FS.
[0037] Regarding drift and self-calibration, to mitigate the impact of temperature and aging on the measurement chain, this step performs a short-term self-calibration during vehicle idling or low-load periods: DC bus voltage The channel performs a Fast Fourier Transform on the bus ripple under no-load conditions to estimate the noise floor. If the noise floor exceeds a threshold, the window averaging time is automatically increased; three-phase current... The shunt and Hall dual-channel circuitry performs zero-point reestimation and gain comparison in the zero-current range; piecewise linear correction is triggered when the deviation exceeds a threshold; the on-state voltage drop sampling value... under constant gate control signal Repeated sampling is performed under narrow pulse conditions, and grid regression is used to eliminate slow drift caused by temperature rise; device junction temperature estimate. The dual-mode thermal observer reverted to the equivalent π-type thermal network mode within a long-term stable range to update the prior static thermal parameters. After verification with 50 sets of train operation data, the above self-calibration process reduced the cross-day measurement repeatability from ±0.8%FS to ±0.3%FS, with no abnormal triggering occurring within a 30-day period.
[0038] Regarding data quality control and metadata management, this step generates quality flags for each sampling window, including saturation detection, sample loss detection, synchronization deviation detection, anti-aliasing margin detection, temperature condition flags, and electromagnetic interference event flags; simultaneously, it also generates operating condition parameters associated with the window. , Record interval numbers and slope information for subsequent feature statistics and model updates at the interval level. All raw and preprocessed data are written to a read-only log using a four-dimensional index of "signal source - timestamp - channel number - topology location". The log includes the acquisition firmware version number, gain matrix version number, and time base checksum to meet the requirements of subsequent traceability and versioned model training.
[0039] It should be noted that this step also includes the following feasible methods: First, introducing perturbation adaptive gating: when the operating condition parameters Automatically suppress diagnostic disturbances when at peak traction or braking speed. Amplitude to maintain traction error The resistance remains consistently below a threshold, thus ensuring observability without introducing perceptible traction fluctuations. In 500 kilometers of test track data, this gating strategy ensures that the equivalent series resistance of the DC link is maintained. Under the premise of no variance degradation, the peak-to-peak value of traction fluctuations is reduced by approximately 38% compared to the fixed perturbation strategy. Secondly, a pre-correction of channel mapping based on prior structural awareness is incorporated: before generating the channel-submodule topological correlation matrix M, submodule-level phase shift compensation is performed on the fixed delay caused by unequal physical wiring lengths to improve the subsequent submodule-level eigenvectors. The comparison between groups A and B shows that this pre-correction can make the spectral sideband energy-based... The sub-module clustering profile coefficient improved by 0.08 to 0.12. Thirdly, a safety-oriented measurement reliability score was implemented: combining quality flags and cross-channel consistency, a measurement reliability score was calculated and compared with the device junction temperature estimate. The uncertainties are used together as weights for subsequent optimization constraints, thereby maintaining the conservatism of decision-making in measurement degradation scenarios.
[0040] In summary, this step achieved a systematic engineering design for traction converter applications in areas such as sensor chain, synchronization, anti-aliasing, calibration, and quality control, laying the foundation for subsequent applications in the frequency band. Internal calculation of equivalent series resistance of DC link Extracting the gate charge integral Duration of shutting down the Miller platform Spectral sideband energy It provides a sampling basis with high signal-to-noise ratio and high temporal consistency, and also provides a basis for the log-likelihood ratio statistic for change point discrimination. Stability estimation and device junction temperature estimation based on dual-mode thermal observer output. The continuity is guaranteed.
[0041] In this embodiment, step S2 will be described in detail. This step involves injecting a small-amplitude diagnostic perturbation into the pulse width modulation link. To stimulate and identify the health-related dynamics of the traction converter, and to constrain traction error. The following will be determined by the diagnostic perturbation. The resulting DC-side response quantity voltage change Synchronously record the changes in current to facilitate subsequent operation in the preset frequency band. Internal estimation of equivalent series resistance of DC link And calculate the spectral sideband energy Lay the foundation for data.
[0042] First, regarding the modulation injection location and signal morphology, the diagnostic perturbation... Superimposed on the modulation ratio by micro-jittering of duty cycle The command channel employs a zero-mean, band-limited, energy-controlled sequence to avoid DC bias. To achieve good identification observability and feasibility, the diagnostic perturbation... It supports two types of modes: one is single-frequency sinusoidal injection, corresponding to the perturbation frequency. Strictly meet carrier frequency The duration is determined by the carrier PWM setting; another type is pseudo-random binary sequence PRBS injection, with a duration of... One carrier cycle, with an upper limit of amplitude. and limited This ensures sufficient identification bandwidth without disrupting modulation linearity. In engineering implementation, single-frequency injection is prioritized for high signal-to-noise ratio identification of the target frequency, while PRBS injection is prioritized for fast scanning and full-band modeling; the two can be mutually exclusive during operating condition switching to reduce the impact on traction comfort.
[0043] Secondly, regarding frequency and amplitude tuning, the diagnostic perturbation... perturbation frequency A carrier-locked loop (CLL) grid selection strategy is adopted, namely... ,in and Use coprime integers to avoid visible beat frequencies; the perturbation frequency The minimum resolution is determined by Decision, typical setup This is to cover the critical frequency band of the component's thermal-electric coupling. To ensure that no perceptible traction torque pulsations are introduced, the diagnostic perturbation... upper limit of amplitude Tuning is performed according to the dual-loop principle of "prior constraints + online feedback": the initial value is given offline using the system rated torque-duty cycle sensitivity curve. The traction error is corrected online using a traction error monitoring loop to constrain the traction error. Always satisfy and leave room for flexibility In one embodiment, A value of 0.006-0.012 can be used. At a Hz setting, ensure that cabin comfort levels do not decrease, while simultaneously maximizing spectral sideband energy. The signal-to-noise ratio is improved by 7-11 dB.
[0044] Furthermore, regarding synchronous recording and demodulation measurements, in order to ensure the voltage change... With the change in current For the diagnostic perturbation The response can be robustly extracted. This step shares a unified Precision Time Protocol (PTP) boundary clock between the acquisition and control ends, and then uses phase-locked reference demodulation plus sliding window averaging to perform amplitude-phase estimation. Specifically, for single-frequency injection, the reference demodulator uses the perturbation frequency... The inner product is performed using orthogonal bases, and the real and imaginary parts are output as estimates of the complex amplitude. For PRBS injection, the matched filter uses a known pseudo-random binary sequence template, and the corresponding generalized correlation peak is the response intensity estimate. To balance real-time performance and variance convergence, the sliding window length is designed as follows: (PRBS) or (Single frequency), with added Hanning weighting to reduce leakage; on a 1 MW traction converter test bench, the voltage change can be measured within a 120 ms observation window using the above method. With the change in current The standard deviations were reduced to 2.1% and 2.7% of the mean, respectively, which was the subsequent... This lays the foundation for controlling the relative uncertainty at 3.5%.
[0045] In terms of constraints and safety, this step establishes a triple protection chain to ensure operational safety and recoverability. The first layer is the gating logic for the traction-braking interval: when operating parameters... Within the maximum traction or maximum braking range and speed Below the threshold At that time, the diagnostic perturbation will be automatically... upper limit of amplitude Reduce to zero to avoid large torque pulsations; when speed Above the threshold At this time, only single-frequency injection is allowed and PRBS injection is disabled to reduce passenger-carrying perception. The second layer is a protection trigger interlock: when the device junction temperature estimate is... Exceeding the warning threshold or DC bus voltage If the ripple index exceeds the limit, immediately freeze the diagnostic perturbation. It also clears the modulation increment; after the interlock condition is released, it recovers exponentially. The third layer is to perform saturation-inverse integration coupling: at the modulator port, the injection loop and the main modulation loop share the limiting-inverse integration unit to avoid the sudden power-on change caused by the accumulation of errors in the integrator during the injection freeze period.
[0046] In terms of data and results, bench and line tests show that, under the same sampling rate conditions, compared to "passive observation" without injection, the diagnostic perturbation method described above is more effective. After that, the energy of the spectral sidebands The measurability has been significantly improved, in kHz , With this combination, the sideband peak signal-to-noise ratio increased from 9.4dB to 19.8dB; the equivalent series resistance of the DC link calculated from this is... of Repeatability improved from ±8.7% to ±3.2%. In a comparison of two groups of aged prototypes, this improvement resulted in earlier visibility of capacitive degradation, equivalent to a reduction of approximately 28% in the deviation of remaining life estimation over a three-month observation period (based on the equivalent series resistance baseline obtained from follow-up disassembly and testing). Simultaneously, when using the aforementioned traction error constraint... After implementing the gate control strategy, the longitudinal acceleration RMS of the passenger compartment did not increase significantly, meeting the comfort evaluation criteria.
[0047] It should be noted that this step also includes the following optional embodiments: First, the "orthogonal multiphase coding" extension of the injection mode: Walsh-Hadamard code sequences that are orthogonal to each other are injected into the three-phase bridge arms respectively, and the responses of each bridge arm are separated by orthogonal correlation within the same observation window, thereby improving the submodule-level feature vector without increasing the total disturbance energy. The extension improves the separability of sub-module clustering, reducing the misclassification rate from 6.3% to 2.4% on a three-level topology prototype. Secondly, a joint frequency transition-residence strategy is employed: "narrowband residency + long window averaging" is used during the stable healthy state, while "multi-point frequency transition" is used during the suspected degradation period. This is achieved by comparing different... The amplitude and phase response are used to distinguish between two mechanisms: "capacitor ESR increase" and "conductor / busbar contact resistance increase"; this strategy improved the AUC of both mechanisms from 0.86 to 0.93 in the fault attribution experiment. Thirdly, online SNR adaptation: based on the aforementioned voltage change... With the change in current The variance is used as a proxy, and the minimum required inverse solution is obtained according to the target SNR. With the shortest dwell window And subject to the traction force error constraint. Together with safety interlocks, it achieves the principle of "injecting only what is needed" with minimal disturbance. Fourth, disturbance energy budgeting and traceability: a disturbance energy budget ledger is established for each operating segment to record the diagnostic micro-disturbances. The frequency, amplitude, dwell time, and gating status, combined with operating condition parameters. , With device junction temperature estimate Generate injection compliance reports to facilitate operational security audits.
[0048] In summary, this step, with its rigorous injection-acquisition-demodulation-constraint-interlock closed-loop design, significantly improves the observability of health-related dynamics and limits the impact of perturbations on traction quality to a measurable range; combined with subsequent steps for spectral sideband energy... DC link equivalent series resistance Duration of shutting down the Miller platform The extraction and statistical inference of features can reveal degradation trends at an earlier stage and shorten the maintenance prediction time window.
[0049] In this embodiment, step S3 will be described in detail. This step is performed within a preset frequency band. Internal DC link equivalent series resistance Calculations are performed, and the gate charge integral is extracted from the aforementioned acquired signals. Duration of shutting down the Miller platform Spectral sideband energy The preset frequency band Used to limit the effective range of subsequent frequency domain estimation to avoid low-frequency operating condition disturbances and high-frequency carrier harmonic pollution; the equivalent series resistance of the DC link The loss level of the DC bus capacitor is reflected; the gate charge integral... From the gate current waveform The duration of the turn-off Miller plateau is obtained by integration over one switching cycle and is used to characterize the device's gate drive and switching behavior; The transient collector-emitter voltage waveform of the switch The plateau segment characteristics are calculated to reflect the junction capacitance and current transfer characteristics during the turn-off process; the spectral sideband energy is... In the set of sideband frequencies The above calculation is used to quantize the carrier frequency. With perturbation frequency The intermodulation response intensity.
[0050] First, regarding the equivalent series resistance of the DC link... The calculation is based on the voltage change. With the change in current The ratio is estimated. To reduce the effects of noise and phase mismatch, the voltage variation is estimated in the frequency domain. With the change in current Bandpass filtering is performed, and the bandwidth is the preset frequency band. And a reference-based coherent demodulation is used to align the two to the same phase reference, and then calculations are performed. When using single-frequency injection, and Take the amplitude of the coherent component; when using pseudo-random binary sequence injection, and The equivalent amplitude corresponding to the peak value of the matched filter is obtained. To suppress random errors, this step performs N independent estimations within a sliding window and takes the weighted median as the final value. The weights are determined based on the signal-to-noise ratio of each estimation. In some embodiments, , , Under a 120 ms observation window and a 10th-order sub-estimation condition, the equivalent series resistance of the DC link is... The three-sigma relative uncertainty is no higher than 3.5%, and the average deviation compared with the disassembly and inspection benchmark is less than 2%.
[0051] Secondly, regarding the gate charge integral... The extraction is first based on the gate control signal. The periodic boundary in the gate current waveform Gating time window for capturing a single switching cycle To compensate for baseline drift and measurement bias, a small reference interval is taken before and after the window to estimate the zero point and perform linear bias correction. Then, the calculation is performed using standard integral form. During the IGBT turn-off phase, to avoid the impact of high-frequency ringing during the Miller plateau on integral stability, a combination of three-point median filtering and a 500 kHz low-pass bandgap is used to optimize the gate current waveform. Bench tests show that when the ambient temperature rises from 25°C to 85°C, the gate charge integral... The mean variation is less than 4%, and the standard deviation is controlled within 2% of the mean, which can be used as a temperature-robust switching behavior characteristic for subsequent state inference.
[0052] Then, regarding the duration of the shut-off Miller platform... The extraction first involves analyzing the transient collector-emitter voltage waveform during switching. By performing the first difference, we obtain The start and end points of the Miller platform are identified using a method combining zero-crossing counting and threshold hysteresis. The specific process is as follows: First, calculate... The following steps are taken: 1) Calculate the local mean and standard deviation of the zero-crossing threshold to exclude small ringing; 2) Find the first zero-crossing point where the platform transitions from a sharp negative change to a steady state as the platform starting point; 3) Find the next zero-crossing point where the platform exits and enters a steady state as the platform ending point; 4) Calculate the time difference between the two points. To improve robustness, a second-order polynomial fitting is performed on the platform segment to smooth out minor ringing, and the extracted data is verified using fitting residual constraints. In the statistics of 5 groups of device samples, each with 2000 shutdown records, the duration of the shutdown Miller plateau was... The repeatability is better than ±5 ns, and it shows a monotonically increasing trend with the increase of device aging cycles (an average increase of 8 to 15 ns per 100 hours of aging), consistent with the junction capacitance changes obtained from disassembly and inspection.
[0053] Next, regarding the spectral sideband energy The calculation uses carrier frequency. Centered, perturbation frequency The set of discrete frequency points with offset To obtain a high-resolution frequency domain estimate, this step involves sampling at the frequency... The following steps involve using integer-period window registration and zero-padding of the sequence to ensure the target frequency falls on a discrete frequency grating. For single-frequency injection, the Goertzel algorithm is recommended to directly calculate the complex spectrum of the target frequency and take the square of the amplitude as the energy. For pseudo-random binary sequence injection, correlation demodulation is first performed using a pseudo-random template, and narrowband energy accumulation is performed at the set of sideband frequencies to form... ,in This represents the complex spectrum at the corresponding frequency. To compensate for window leakage and amplitude bias, an equivalent noise bandwidth normalization factor is used to correct the energy. In the line operation data, when At that time, the spectral sideband energy The median signal-to-noise ratio is improved by approximately 10.4 dB in the no-injection condition, and this is in contrast to the equivalent series resistance of the DC link. It shows a significant positive correlation (Pearson correlation coefficient 0.72), which can be used as a mutual verification feature of degradation intensity to improve the discriminative power of failure probability estimation.
[0054] To ensure the physical consistency and statistical validity of the above characteristics, this step provides the following data quality control and uncertainty assessment strategies. First, by integrating the gate charge... Duration of the shut-off Miller platform A robust regression is performed on the joint distribution to remove outliers caused by abnormal driving pulses or sampling saturation; secondly, by adjusting the equivalent series resistance of the DC link... The repeated measurement sequences are subjected to variance decomposition to quantify the three types of uncertainty caused by injection amplitude jitter, demodulation window length variation, and background condition disturbance, and these are represented as observation noise covariance in subsequent variational inference; third, the spectral sideband energy is analyzed... With the carrier frequency The symmetry of the side band structure is checked, if... and If the energy difference exceeds the threshold, resampling or a change of the window function is triggered to avoid bias accumulation.
[0055] It should be noted that this step also includes the following optional implementation methods: First, introducing adaptive frequency band selection: automatically updating the preset frequency band based on the real-time estimated injection signal-to-noise ratio and background spectral line positions. The upper and lower limits make and It remains in the passband plateau region, thereby suppressing window leakage from affecting the spectral sideband energy. The system bias. Secondly, coherent multi-window spectrum estimation is adopted: energy estimates of multiple orthogonal windows are calculated in parallel within the same observation window and their unbiased synthesis is taken to reduce the variance of the frequency domain estimation; this method is applied to the equivalent series resistance of the DC link. The frequency domain ratio calculation can further reduce the relative uncertainty to approximately 2.5%. Thirdly, a joint fitting with physical consistency constraints is introduced: the gate charge is integrated... The duration of the shut-off Miller platform With the spectral sideband energy As joint observations of the same switching unit, weighted least squares are used to solve a set of "degradation intensity-temperature-current" coupling parameters within the same time window to improve characteristic stability under temperature drift conditions. Fourth, a sideband energy ratio index is constructed: defining... When the system symmetry is compromised or the inconsistency of the bridge arm components increases, The magnitude of the deviation from 1 can be used to detect imbalance degradation early; in the two-submodule inconsistency experiment, the magnitude of the deviation from 1 can be used to detect imbalance degradation early; in the two-submodule inconsistency experiment, the AUC reaches 0.91.
[0056] Through the above process and improvements, the equivalent series resistance of the DC link The gate charge integral The duration of the shut-off Miller platform With the spectral sideband energy All of these can be stably extracted under a unified time base and statistical control, and compared with the voltage change obtained in the previous steps. The change in current This, along with subsequent sub-module-level features based on the topological association matrix, forms a consistent data loop.
[0057] In this embodiment, step S4 will be described in detail. The core objective of this step is to generate a directed graph of three-phase-bridge arm-submodule based on the topology configuration file and construct the channel-submodule topology correlation matrix M. Then, the multi-source signals acquired in the preceding steps are mapped into submodule-level feature vectors. The topology configuration file defines the electrical structure and measurement point layout of the traction converter, including three-phase numbering, the structure of each bridge arm, the number and sequence of submodules, the relationship between bypass and switching switches, the correspondence table between measurement channels and physical nodes, and the channel numbers of sampling and control boards. The topology configuration file is stored in a key-value structure and includes a version number and checksum to ensure the uniqueness and traceability of software parsing.
[0058] First, a directed graph of three phases, bridge arms, and submodules is established based on the topology configuration file. .in, It is a set of nodes, including phase nodes, upper arm bus nodes, lower arm bus nodes, DC bus positive and negative pole nodes, and port nodes of each submodule; Let be a set of directed edges, representing the directed equivalent path of current from the power source to the load. For ease of algorithm operation, we will use a directed graph... Calculate the adjacency matrix A and the node-edge incidence matrix B. The elements of A... This indicates the existence of a self-node. Pointing to node A directed connection is defined if B is true, otherwise zero; elements of B or Indicates the first The start and end points of a directed edge are relative to the first... The incident symbols of each node. This graph structure will be used to link the measurements of the channel layer with the health status of the submodule layer. For different topologies (two-level, three-level midpoint clamped NPC, active midpoint clamped ANPC, modular multilevel MMC), the topology generator automatically instantiates nodes and edges according to the "arm type - number of submodules - bypass rules" in the configuration file; for example, in the modular multilevel MMC topology, if each arm has N submodules and adopts a series structure, the submodule port nodes are linearly connected to the arm bus nodes in sequence {1,...,N}, and bypass edges are added to represent the current path in the fault or deactivated state.
[0059] Secondly, construct the channel-submodule topology correlation matrix M. Rows in M correspond to acquisition channels, columns correspond to submodules, and elements... Indicates the first The acquisition channel for the first... The "observable contribution" of each submodule. For measurement points directly mounted on the device or submodule (e.g., switching transient collector-emitter voltage waveforms)... Gate current waveform On-state voltage drop sampling value The mapping between "test point - device ID" in the configuration file can directly enable the corresponding... Other columns are zero; for aggregated signals such as phase current and bridge arm current (e.g., three-phase current), the values are zero. The circuit structure is used as a priori to distribute its energy and phase to the submodule columns on the respective bridge arm links according to the current splitting relationship, forming a sparsely distributed weight vector. To ensure scale consistency, column normalization is performed on each row to obtain a column random matrix form, making... For all Established. Considering the fixed delay caused by the unequal physical distances and wiring lengths of different channels, this step applies submodule-level phase shift compensation before constructing matrix M. The compensation amount comes from the metadata "route length—medium propagation speed—board clock offset" in the topology configuration file. After compensation, the phase alignment information required for coherence index calculation is written into the auxiliary dictionary of M for subsequent unified use. To improve robustness, when a channel fails or degrades, the corresponding row will be downweighted and adaptively filled with a weight template of "adjacent alternative channels," thereby maintaining the sparsity and observability of M. This process can be regarded as constraining... The minimum deviation reconstruction problem is defined as follows, with the objective of minimizing the topology consistency loss and time alignment residual.
[0060] Next, complete the mapping from channel-level features to sub-module-level features. Let the channel-level feature vector be... ,in This indicates a standardized extraction operator for waveform features (such as mean, peak, coherent components, or band energy), where all components are aligned within the same time window and normalized according to dimensions. After copying and expanding along the channel dimension, multiply it on the left by the channel-submodule topological association matrix M to obtain the submodule-level feature stacking vector. In the formula This is the feature concatenation vector after channel expansion. For Kronecker product, For the first The feature vectors of each submodule. To improve the consistency of submodule-level features, this step... Two types of constraints are applied: one is the energy conservation constraint, which ensures that the sum of the current characteristics distributed from the phase current or bridge arm current to each submodule is statistically consistent with the original channel characteristics; the other is the topological smoothing constraint, which introduces the Graph Laplace. ,by The penalty term suppresses non-physical abrupt changes in adjacent sub-modules on the same arm, thereby enhancing the physical interpretability and noise resistance of features.
[0061] Subsequently, to address the topology time-varying conditions caused by bypassing, deactivation, and thermal redundancy, this step constructs a time-varying channel-submodule topology correlation matrix. When the gate control signal When the switch status indicator shows that a submodule is in bypass mode, the weight of the direct measurement point in the corresponding column is maintained, while the aggregate weight allocated through the bridge arm current is cleared to zero or scaled according to the "bypass equivalent resistance" model; when the bridge arm is deactivated and triggered... or When the health management instruction is executed, it is based on the bridge arm switching logic. The current and voltage paths are redistributed in the middle. To avoid jitter caused by frequent switching, The update adopts a hysteresis strategy, and the update is only implemented when the bypass state continues for more than the minimum dwell time; and the historical data in the window is smoothed by exponential weight coefficients during the update to ensure the stability of subsequent statistical inference.
[0062] To achieve traceability and consistency verification, this step sets up three checks for topology generation and matrix construction. The first is a static consistency check: checking if the weight sum of each row is 1, and checking if the non-zero elements of each column are consistent with the prior structure of which channels the submodule can be observed. The second is a dynamic consistency check: within a time window, for... The energy distribution of phase current and bridge arm current is compared with the actual measured values. If the deviation exceeds the threshold, the system rolls back to the previous version and triggers route delay recalibration. The third step is version and hash verification: verifying the topology configuration file, the generated directed graph, and the channel-submodule topology association matrix. Calculate the hash and write it to the log so that each feature mapping corresponds to a unique structure version.
[0063] In the combined bench and line experiment, the channel-submodule topological correlation matrix M was used and smoothed using a graph Laplace, based on the spectral sideband energy. Integral with gate charge The submodule clustering purity is significantly improved compared to black-box mapping without structural priors, with an average improvement of 0.11 in the silhouette coefficient. The Top-1 accuracy of submodule-level fault location is improved from 78% to 93%, and the Top-3 accuracy reaches 98%. In time-varying scenarios involving bypassing and deactivation, the time-varying channel-submodule topological association matrix... Preserves submodule-level feature vectors The time consistency ensures that the performance degradation of the degradation coefficient estimation model trained in this way is less than 3% in the training-deployment domain offset test.
[0064] It should be noted that this step also includes the following optional embodiments: First, introducing structure-aware group sparsity weight learning: on small-scale labeled data, group sparsity and... The system employs a joint regularization approach, fine-tuning M under constraints of row summation of 1 and sparsity upper limits to maximize the separability of submodule features even with inconsistencies in wiring and measurement point deviations. Secondly, it utilizes spectral embedding for topology self-checking: A is eigenvalued, and the angle between the eigenvectors and historical benchmarks is monitored. When the angle exceeds a threshold, a warning is issued indicating potential inconsistencies between the topology configuration file and the actual wiring, automatically generating correction suggestions. Thirdly, it establishes insertion duty cycle statistics for MMC: Effective insertion duty cycles for each submodule are statistically analyzed within the same observation window, serving as... The column scale factor participates in feature normalization, weakening the systematic bias caused by changes in modulation strategy on the feature distribution of sub-modules. Fourth, the "measurement point replacement optimization" strategy is implemented: when a direct measurement point temporarily fails, the channel with the largest weight is automatically selected from the set of alternative channels to fill the gap based on the topology path length and historical consistency score, thereby avoiding the interruption of the overall feature pipeline.
[0065] Through the above steps, the directed graph of the three-phase-bridge arm-submodule and the topological association matrix of the channel-submodule can be automatically generated in engineering and their correctness and traceability can be ensured through multiple consistency checks. Submodule-level feature vectors... It can be stably constructed under a unified time base and structural prior constraints, providing high-quality and interpretable submodule layer inputs for parameter identification and variational inference of subsequent degenerate state space models, as well as optimization decisions based on safety barrier functions.
[0066] In this embodiment, step S5 will be described in detail. This step establishes a degenerate state-space model. Among them, the state vector This involves aggregating low-dimensional "latent degradation" related to health, represented by submodule-level feature vectors. The robust principal components are obtained through structural prior compression and used to characterize the slow degradation trend and operational sensitivity of devices and sub-modules; input vector Operating condition parameters The structure is used to characterize the exogenous influence of traction / braking mode and speed on the system response; the output vector It is composed of stacked observational features used for diagnosis and prediction, including the equivalent series resistance of the DC link. Gate charge integral Duration of shutting down the Miller platform Spectral sideband energy On-state voltage drop sampling value Device junction temperature estimate The combination of components (normalized by dimensions and aligned to the same time window); system matrices A, B, and C respectively describe the temporal evolution of degradation, the linear influence of operating condition excitation on degradation evolution, and the mapping of degradation to observed features; process noise. With observation noise It is regarded as a Gaussian disturbance with zero mean and covariances Q and R, respectively, and is used to model unmodeled dynamics, measurement errors and environmental disturbances.
[0067] To ensure feasibility and robustness, the degenerate state-space model is established using a two-stage strategy of "offline structure identification + online fine-tuning". In the offline stage, the topology configuration file and the channel-submodule topology association matrix M are used as structural priors to identify submodule-level feature vectors. Perform robust principal component analysis on a multi-condition dataset, selecting interpretable top-valued data. The principal components are used as the initial state vector. The composition dimension is determined, and non-physical jumps between adjacent sub-modules in the same bridge arm are suppressed using graph Laplacian regularization; subsequently, strip group sparsity and... Regular least squares estimation of A, B, C, constraints (Spectral radius not greater than 1) to ensure asymptotic stability, constraining the block sparse structure of C to preserve the physical correspondence between "features and sub-modules"; finally, the covariance Q and R are estimated using the maximum likelihood method or the expectation-maximization strategy. In the online phase, under the premise of fixed structure and regularization terms, recursive least squares fine-tuning is performed on the small drift parameters of A and B using a sliding time window, and R is adaptively updated through observation residuals to cope with slow drift caused by temperature, load and aging.
[0068] To promptly identify abrupt changes or rate alterations in the degradation mechanism, this step constructs change point indicator variables. and the log-likelihood ratio statistic for the change point And set a threshold for detecting change points. Specifically, a generalized likelihood ratio test is established using two alternative hypotheses: : in the past window Within this period, the parameters of the degenerate state-space model remain unchanged; The parameter mutation or process noise covariance Q increases abruptly at an unknown time within the window. An innovative sequence is obtained using a Kalman filter. and its covariance Calculate the normalized innovation energy at each moment. And form a generalized likelihood ratio within the sliding window. ,in This is the expected value under the assumption of no variables. If Then place This triggers online thermal parameter identification and model covariance reestimation; simultaneously, robust weights are applied to data near the change point to avoid instantaneous parameter oscillations. Threshold The specified false alarm probability With window length Pre-calibration using the Monte Carlo method, and taking data from the bench setup. , At that time, the average detection delay was less than 120 ms, and the false alarm rate was controlled within 1.1%.
[0069] To ensure numerical stability and engineering consistency, the degenerate state-space model introduces four types of engineering constraints. First, stability constraints: through the Shure decomposition of A and pole projection, the eigenvalues near the unit circle are shrunk to the radius. In some embodiments, A value of 0.02 is used to avoid boundary instability caused by online fine-tuning. Second, nonnegativity and monotonic priors: nonnegativity constraints are applied to the state components that directly reflect the degradation intensity, and small constraints are applied to their first-order differences in adjacent time windows. Regularization is used to conform to the general law of "slow and monotonic" degradation. Third, structural sparsity constraints: the sparse structure derived from the topology is preserved on the "characteristic-state" mapping of C, making the DC link equivalent to a series resistance. Spectral sideband energy More dependent on the state components related to the capacitor-DC link, gate charge integral Duration of shutting down the Miller platform It relies more on the state components related to the switching unit. Fourth, flexible constraints on operating conditions: in the two columns of B (corresponding to operating condition parameters) Apply an upper bound to the norm to prevent changes in operating conditions from being incorrectly interpreted as degenerative changes by the model.
[0070] Regarding its connection with subsequent steps, the degenerate state-space model outputs two key types of information: one is the prior distribution and innovation sequence that can be used for mean field variational inference, and the other is an instantaneous consistency index for thermal parameter re-identification and safety barrier function evaluation. Specifically, the state estimate is based on Kalman filtering and smoothing. and covariance This serves as the initialization for subsequent mean field variational inference; simultaneously, it sets the change point indicator variable. Log-likelihood ratio statistic with change point discriminant Provided to online thermal observers to trigger equivalent The thermal network parameters are identified and updated to maintain the device junction temperature estimate. Consistency. For the calculation of the safety barrier function, the degenerate state-space model is obtained through... The predicted values help evaluate the consistency between the "predicted remaining lifetime and current observations" and are fed back to subsequent optimization solutions.
[0071] In the joint bench and line experiments, using the above modeling process, the degraded state-space model can remain stable under conditions of multi-condition switching and bypass / deactivation, and innovate statistics. The distribution is consistent with the baseline under the assumption of no invariant point; artificial degradation is introduced (increasing the equivalent series resistance of the DC link). In an experiment with approximately 20% of the data points, the change-point discriminant log-likelihood ratio statistic... Stably crosses the threshold within 90–140 ms No false alarms were detected in the false alarm experiment (random perturbation without changing the device or operating conditions). Offline cross-validation showed that, with the addition of structural sparsity and non-negative-monotonic priors, the mean square error of one-step prediction of key observations was reduced by 18%-31% compared to the black-box state-space baseline without priors.
[0072] It should be noted that this step also includes the following optional implementation methods: First, introducing block diagonal state evolution with graph Laplace regularity: the state vector Taking bridge arms or submodule clusters as blocks, let A have a block diagonal plus sparse nearest neighbor coupling structure, and apply a graph Laplace penalty to the cross-block coupling terms, thereby directly embedding the topological prior into the time evolution matrix. Secondly, temperature coupling process noise adaptation: let the process noise covariance Q be the estimated value of the device junction temperature. The affine function allows for stronger diffusion in the state evolution at high temperatures, consistent with the physical mechanism of "accelerated degradation due to thermal stress." Thirdly, robust statistics of innovation energy: Normalized innovation energy is constructed by replacing the quadratic loss with Huber loss, reducing the impact of single outlier sampling on the log-likelihood ratio statistic for identifying change points. The impact. Fourth, explicit jump modeling of operating condition switching: in the input vector The state codes of the merged finite state machines are used to form an equivalent representation of the switching linear model, enabling the model to automatically switch to the corresponding state at the instant of traction / braking switching. Subsets improve the accuracy of transient prediction and the specificity of change point localization.
[0073] In summary, the degradation state-space model, with structural priors, stability, and physical consistency as its core, forms a feasible modeling and updating mechanism from offline to online. Combined with the change point indicator variable and the change point discriminant log-likelihood ratio statistic, it can reliably respond to changes in degradation mechanisms at an early stage. Quantitative results validated on bench and line show that this modeling framework meets the requirements for subsequent probability and lifetime joint inference and safety barrier optimization decision-making in terms of prediction accuracy, false alarm control, and engineering stability.
[0074] In this embodiment, step S6 will be described in detail. This step involves the degenerate state-space model. Perform mean-field variational inference and output a fault occurrence probability vector. Remaining life estimate and remaining lifetime uncertainty Among them, the failure occurrence probability vector For each fault mode at time Conditional probability components, remaining lifetime estimate From time The statistically expected time until the failure criterion is reached, and the remaining lifetime uncertainty. This represents the corresponding standard deviation or equivalent confidence interval scale. To simultaneously address gradual degradation and mechanistic abrupt changes, the latent variables in this step include the degradation state. With variable point indicator variable It employs an inference method that combines a "time sliding window + online recursion".
[0075] First, we construct the factorized form of the mean field. Let the posterior approximation be... ,in For Gaussian time series approximation, and They represent the times respectively. The state mean and covariance; It is an approximation of the binomial distribution. Let's consider the probability of the change point. The objective is to maximize the lower bound of the evidence. Under the premise of fixing the system matrices A, B, C and the noise covariance Q, R, Coordinate ascent optimization is performed. Secondly, a variational information filtering-smoothing update formula is given. This is applied to the Gaussian time series approximation. Kalman filtering in informational form and RTS smoothing are used to obtain numerically stable mean and covariance updates; in each prediction and correction step, the expected value of the change point indicator variable is approximated. The mapping is "augmented process noise", that is, let ,in This is a small diffusion coefficient, used to improve the model's tolerance to mutations at suspected change points. From this, we obtain the recursion: predict: , Correction: , .
[0076] Then smoothed using RTS. As a moment The final state is approximated posteriorly. This approximation is made for the binomial distribution. Using innovation statistics The log-likelihood ratio increment is constructed from the deviation from the mean of the unchanging baseline, and then updated. ,in , , The gain coefficient is calibrated offline; it is used to suppress jitter. This is achieved through time-domain hysteresis and exponential smoothing.
[0077] Then, generate a failure probability vector. Let the log-probability vector for fault classification be... ,in , The linear mapping parameters are obtained through offline learning and are defined according to softmax. .
[0078] To maintain physical interpretability, The rows are subject to a set of sparse constraints based on the feature-state structure prior, making the state components related to capacitor degradation more influential on the "DC link type fault" components, and the state components related to switching unit degradation more influential on the "device type fault" components; simultaneously, for By incorporating linear covariates of temperature and current, we can avoid misjudging operating condition drift as a change in fault probability.
[0079] Next, an estimate of the remaining lifetime is generated. With remaining lifetime uncertainty A degradation-to-failure model based on discrete-time hazard rate is adopted: Let the hazard rate... , ,in From time The mean of the predicted states obtained by forward propagation of the degenerate state-space model. , These are the coefficients from offline learning. The survival probability is obtained from this. The remaining life estimate is taken as the expected value. The remaining lifetime uncertainty is assessed by combining first-order error propagation with Monte Carlo sampling: The sample forward prediction is obtained Take the sample standard deviation as In online implementation, to meet the time budget , Selected as satisfied The smallest integer, Pick .
[0080] Regarding numerical stability and engineering robustness, this step employs three measures. First, maintaining positive covariance definiteness: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] and Apply Jolesi decomposition and spectral clipping to prevent nondefiniteness caused by numerical rounding; perform diagonal loading when the condition number exceeds the threshold. , For adaptive small constants. Second, anomaly observation suppression: for innovation. We employ Huber loss-equivalent weight pruning and incorporate anomaly levels into the weighting. Gain adjustment to avoid spurious changes caused by single-point anomalies. Third, online sliding window and forgetting: in a length of... Recursive inference within a time window, applying a forgetting factor to remote data. For example, take Achieve continuous adaptation to slow drift.
[0081] In the combined bench and line test, compared with the baseline of "variableless fusion", the fault occurrence probability vector after using the mean field variational inference is... The Top-1 classification accuracy improved from 86.7% to 92.4%, while the false positive rate decreased by 31%; the remaining life estimate calculated using the aforementioned hazard rate model... The median absolute error relative to the disassembly and inspection benchmark decreased from 14.2% to 8.9%. In time-varying scenarios involving bypass and deactivation events, the remaining lifetime uncertainty... The Pearson correlation with the actual error reaches 0.78, demonstrating the confidence level of the uncertainty. Furthermore, in the artificial degradation experiment (increasing the equivalent series resistance of the DC link)... In 20%, the probability of the change point is... It rose to 0.82 within 120ms, bringing about... Adaptive diffusion, thereby enabling and It converges to the new degenerate state more quickly.
[0082] It should be noted that this step also includes the following optional implementation methods: First, a variational family of structural a priori constraints: in Introducing graph Laplace regularization into the covariance structure makes the posterior state of adjacent submodules in the same bridge arm smoother and weakens non-physical jumps; this improvement will improve performance under high-noise conditions. The variance is reduced by approximately 12%. Secondly, heteroscedastic observation modeling: Let the observation noise covariance R be the estimated value of the device junction temperature. The function of current amplitude reflects the increased measurement noise under high temperature and high current conditions, which can further reduce the false alarm rate in experiments. Left and right. Third, low-rank-sparse joint prior: for and Introducing low-rank decomposition and Sparse regularization improves generalization across operating conditions and reduces the risk of overfitting; fourth, Bayesian model averaging: weighted averaging by marginal likelihood between two different (A,B,C) structure candidates alleviates the systematic bias caused by model mismatch.
[0083] Through the above process, the failure probability vector The remaining life estimate With the remaining lifetime uncertainty It can be stably generated within a unified variational inference framework and maintains temporal and structural consistency with the output of preceding steps.
[0084] In this embodiment, step S7 will be described in detail. This step involves the fault occurrence probability vector. Implement Dirichlet temperature scaling calibration, temperature scaling parameters and estimates of remaining lifespan. Implement univariate order-preserving regression calibration and monotonic calibration mapping. Output a calibrated fault occurrence probability vector Compared with the calibrated remaining lifetime estimate The goal of this step is to align the probability and lifetime quantization results from the preceding inferences to the true frequency and measured lifetime distribution, thereby ensuring the confidence and cross-condition robustness of the inputs used in subsequent safety barrier functions and optimization decisions.
[0085] First, regarding the fault occurrence probability vector The Dirichlet temperature scaling calibration takes the following form. Let the uncalibrated log-probability vector be... The mapping relationship is , .
[0086] The temperature scaling parameter Controlling the "sharpness" of the probability distribution, when When the distribution becomes smoother, This makes the distribution sharper. To enhance frequency matching under multi-class conditions, this step uses Dirichlet likelihood as the fitting objective, aligning the multivariate distribution frequencies of the true classes within a batch with the temperature-scaled average probability vector. Specifically, this involves minimizing the weighted sum of the negative log-likelihood and the expected calibration error on the validation dataset. ,in Here, NLL is the weight, NLL is the negative log-likelihood, and ECE is the expected calibration error, all of which are related to the model parameters under fixed conditions. Univariate optimization was performed. To avoid overfitting, nested cross-validation was used in this step to determine... With regularization terms, and after each update, the upper and lower bounds of the parameters are scaled by temperature. Crop the data, for example, to [0.5, 5]. During the online phase, adjust the temperature scaling parameter using time window increments. We perform small-step projection gradient updates to adapt to the classification confidence bias caused by changes in ambient temperature and load, while introducing a hysteresis threshold to prevent frequent jitter.
[0087] Secondly, regarding the estimated remaining lifetime... Establishing a monotonic calibration mapping using univariate order-preserving regression calibration. The calibration data consists of paired prediction and actual data from historical samples, with the predictions derived from the remaining lifetime estimate in the preceding inference. The actual lifetime is the observation lifetime followed up to failure or the truncated correction lifetime. To ensure physical consistency and interpretability, the monotonic calibration mapping... The function is constrained to be non-decreasing to maintain the monotonic relationship that "a larger predicted lifetime corresponds to a relatively large actual lifetime". This step uses a pooling adjacency violation algorithm to solve the order-preserving regression problem, obtaining a piecewise constant or piecewise linear form. Meanwhile, extrapolation constraints are added at both ends to prevent non-physical oscillations outside the sample domain. In the online phase, calibration samples are updated using a sliding window, and the monotonic calibration mapping is adjusted. Perform gentle updates; when there are insufficient new samples within the window, reuse the previous stable solution to ensure the calibrated remaining lifetime estimate. Temporal continuity.
[0088] Then, a linkage between quality control and drift detection is established. To ensure that the Dirichlet temperature scaling calibration and the anisotropic univariate ordinal regression calibration remain effective during distribution drift, this step uses the Population Stability Index (PSI) to monitor the distribution difference between the verification data and the baseline data. When the PSI does not exceed the threshold, a conventional incremental update is used; when the PSI exceeds the threshold and the log-likelihood ratio statistic of the change point discriminant is used... Simultaneously, when the temperature increases, the "Quick Calibration" process is initiated: the temperature scaling parameters... The monotonic calibration mapping is updated with larger step sizes. A shorter sliding window is used for re-estimation, while simultaneously triggering extrapolation protection to ensure that the calibration result does not exhibit abrupt changes. To avoid miscalibration caused by a single anomaly, this step performs a fault occurrence probability vector calculation on the input before both types of calibration. Compared with remaining life estimate Implement robust pruning, removing outliers that exceed the quantile threshold or assigning them lower weights.
[0089] Regarding the connection with subsequent steps, the calibrated fault occurrence probability vector Compared with the calibrated remaining lifetime estimate This will serve as direct input to the safety barrier function and the optimization solution. The lifetime constraint terms in the safety barrier function are used... And simultaneously read the variance of the calibration residual as the uncertainty amplification factor; the trigger threshold for probability is determined by... The confidence interval is determined to achieve risk control that allows "confidence to adapt over time".
[0090] In combined bench and line experiments, this step achieved significant improvements in multiple metrics compared to the uncalibrated baseline. Taking a dataset with eight fault modes as an example, the Dirichlet temperature scaling calibration reduced the expected calibration error from 4.7% to 1.6%, decreased the negative log-likelihood by approximately 12%, and maintained or slightly increased the area under the receiver operating characteristic (AUC), indicating that the discriminative ability was not impaired while the calibration performance was significantly improved. The univariate ordinal regression-preserving calibration reduced the median absolute percentage error from 14.2% to 9.1%, and significantly improved the censored log-likelihood. In scenarios involving bypass and deactivation events, the correlation coefficient between the calibrated remaining lifetime uncertainty and the true error reached 0.76, indicating that the uncertainty has strong credibility. Online long-run experiments showed that in the steady-state phase below the PSI threshold, the temperature scaling parameter... The average update rate is less than 0.03 per hour; when there is a disturbance caused by a sudden increase in ambient temperature and a change in operating conditions, the rapid calibration process can restore the expected calibration error to within the threshold within 3 to 5 sliding windows.
[0091] It should be noted that this step also includes the following optional implementation methods: First, category-conditional temperature scaling: Let the temperature scaling parameter... The first method is a function of failure modes, which learns independently on modes with sufficient sample size, while modes with insufficient sample size share a group temperature, thus refining the calibration without significantly increasing the degrees of freedom. The second method is condition-aware calibration: operating condition parameters are used... , As a covariate, the temperature scaling parameter With the monotonic calibration mapping The selection of segmented or hybrid expert strategies allows for calibration within different traction and braking ranges to better reflect actual distributions. Thirdly, truncated consistency correction: for truncated data in the lifespan sample that has not yet failed, partial likelihood from survival analysis is introduced to jointly estimate the monotonic calibration mapping. This improves the reliability of the calibrated remaining lifetime estimate in the early stages. Fourth, uncertainty consistency regularization: A penalty term for "consistency between the predicted distribution variance and the empirical residual variance" is added to the calibration objective, aligning the uncertainty estimate with the actual error, making it easier to incorporate the uncertainty into the safety barrier constraint later.
[0092] Through the above process, the Dirichlet temperature scaling calibration and the anisotropic univariate order-preserving regression calibration work together on a unified data pipeline, making the fault occurrence probability vector... Compared with the estimated remaining lifetime It possesses good calibrability, interpretability, and robustness across operating conditions.
[0093] Step S8 will be described in detail in this embodiment. This step constructs the safety barrier function. And in time budget The internal system uses quadratic programming to generate health management instruction variables. Among them, the safety barrier function The safety margin for constrained peak current is defined as follows: Peak current setting With current upper limit These originate from the real-time control loop and the device's rated boundary, respectively; safety barrier function. Constraint temperature margin, defined as Among them, the upper limit of junction temperature From device specifications, window mean With window standard deviation Estimated value of device junction temperature The safety factor was obtained from statistics within the most recent window. Reflects uncertainty amplification strategies; safety barrier function The lower limit of the constraint lifetime is defined as follows: Among them, the calibrated remaining lifetime estimate Anisotropic univariate order-preserving regression calibration from step S7, stop threshold The acceptable minimum lifetime bound is given by the operation and maintenance strategy. To achieve the minimum cost trade-off between control commands and safety constraints, this step solves the following quadratic programming problem in each control slot: Wherein, objective function Through weight parameters , , A weighted trade-off between "performance preservation" and "deactivation cost" is achieved through modulation ratio. Defined by the modulation link, To fine-tune the increment, Both the fine-tuning amount set for the peak current and the principle of "few actions" are achieved by using the "smallest possible" secondary penalty to activate the switching quantity. Used to isolate bridge arms or sub-modules when necessary.
[0094] To map constraints from the "state space" to the "executable increment," this step introduces first-order sensitivity linearization into the constraints: for and about , Local linear approximation , ,coefficient Obtained from identification regression within the nearest two or three control cycles. To incorporate the deactivation action into convex optimization, this step uses a "continuous relaxation and reprojection" method to process the deactivation switch quantity: Let It participates in the optimization as a continuous variable during the solution phase, and is binarized with a threshold η after the solution is obtained. The threshold η is, for example, 0.5; when the field controller supports a mixed-integer solver, this relaxation can be replaced with mixed-integer quadratic programming to obtain a strict binary decision. Consider... The statistical uncertainty, this step will determine the remaining lifetime uncertainty. Incorporating robustness margin and employing conservative replacement ,coefficient This corresponds to a target confidence level (e.g., 1.96 corresponds to approximately 95%). Similarly, to avoid the probability trigger threshold being influenced by overconfidence, a calibrated fault occurrence probability vector is used. Using risk weights, Dynamic adjustment ,in For the maximum class probability, The degree to which the weight of "high-risk tends to be activated" is increased.
[0095] To ensure the time budget is in place Once an internally stable solution is obtained, this step employs a two-stage strategy of "prediction-feasibility". In the first stage, a warm start is performed based on the Lagrange multipliers and Karoš-Kuhn-Tucker residuals from the previous cycle, using the interior-point method or incremental sequential quadratic programming to solve the main problem; if in If the internal problem has not yet converged, then switch to the feasibility recovery subproblem of the second stage: Among them, slack variables Allowing for limited violation of constraints, This is the penalty coefficient; if slack occurs, the controller operates on the principle of "minimum impact," adjusting only within the necessary range. and set =1 triggers deactivation. To prevent instruction jitter and frequent deactivation, a smoothing term is added to the objective function. ,in For the decision of the previous cycle, For time smoothing weights.
[0096] In terms of numerical stability and engineering robustness, this step establishes three lines of defense. First, soft-constraint barriering: each safety barrier function is embedded into the objective function through a logarithmic barrier. , For decreasing barrier parameters, A small positive number ensures a strong "away" drive near the feasible region boundary; this term can be omitted when using an explicit constraint solver. Secondly, the decoupling-recoupling heuristic: first, in a fixed... Solving continuous variable subproblems under certain conditions, if the optimal value still cannot be satisfied... Then enumerate the candidate deactivation actions in the small set. Thirdly, sensitivity protection: [This is followed by a seemingly unrelated sentence about recalculation and sensitivity protection.] The identification uncertainty is set with an upper bound; the safety factor is automatically increased when the uncertainty exceeds the upper bound. With robustness coefficient The level of conservatism has been raised to a level that guarantees safety.
[0097] Regarding the connection with preceding and following steps, the calibrated remaining lifetime estimate The remaining lifetime uncertainty The calibrated fault occurrence probability vector With the estimated junction temperature of the device Common entry into the safety barrier function and weight scheduling; when the change point is discriminated against, the log-likelihood ratio statistic... When the population stability index (PSI) indicates distribution drift, this step automatically increases the conservatism (increases the population stability index). and And, if necessary, shorten the decision window of "continuous relaxation reprojection" to more quickly transition to deactivation protection. If subsequent steps trigger deactivation, causing a change in the channel-submodule topology correlation matrix M, then the time-varying channel-submodule topology correlation matrix... It will be updated in the next cycle to ensure consistency between structure and control.
[0098] In the combined bench and line tests, compared to the baseline strategy of "threshold alarm + manual derating", this step reduced the peak junction temperature overshoot from 6.8% to 1.9% under the same operating conditions, and reduced the number of emergency shutdowns triggered by overload by approximately 42%. In two manual degradation scenarios (DC link equivalent series resistance...), The 15% increase is related to the aging of the switching unit causing a longer shutdown Miller plateau duration. In the extended 20 ns), Within 5 control cycles after entering the low-lifetime region, the quadratic programming provides... Able to use safety barrier functions At the same time, the non-negative domain was pushed back, and the longitudinal acceleration RMS of the passenger compartment did not increase significantly; in the high temperature and high load disturbance test, after adopting "continuous relaxation reprojection", the false trigger rate of deactivation decision was less than 0.7%.
[0099] It should be noted that this step also includes the following optional implementation methods: First, opportunity constraint generalization: extending the safety barrier function from deterministic constraints to opportunity constraints. ,Will and First, it uses joint calibration with temperature fluctuations, transforming chance constraints into a deterministic half-space through Chebyshev or Gaussian approximations to further reflect the uncertainty-driven safety boundary. Second, it embeds risk metrics: introducing a conditional risk (CVaR) term into the objective function, giving higher weight to the tail risks of lifetime and temperature, and improving robustness in extreme scenarios. Third, it hierarchically structures multiple objectives: placing "maintaining traction performance" as a primary objective and "de-expansion and deactivation" as a secondary objective, through... - Constraint methods or hierarchical weight automatic scheduling implement a hierarchical strategy of "fine-tuning if possible, and reactivating only if fine-tuning is insufficient." Fourth, solver-hardware collaboration: This involves integrating the first-order sensitivity coefficients... The previous optimal solution is sent as a "hot start package" to the real-time solver, and combined with the pre-generated Karoš-Kuhn-Tucker cone linearization template, in The output should guarantee a robust and feasible solution within the given budget.
[0100] Through the above design, the security barrier function Under the unified data and structural priors, quadratic programming forms a closed-loop control, outputting the health management instruction variables. It can minimize performance loss within strict safety constraints, while maintaining the feasibility and interpretability of the solution in scenarios of distribution drift and abrupt degradation.
[0101] In this embodiment, step S9 will be described in detail. In this step, the traction converter control system will, according to the health management instructions obtained in step S8, including the modulation ratio... Peak current and security control disabled flag The health management instructions are issued to the traction converter control unit, and within the specified time budget. Execute internally.
[0102] First, the issuance of health management instructions needs to ensure the system's response speed and accuracy, especially in the face of emergency failures, where instruction execution must be real-time. Modulation ratio The traction converter adjusts the ratio between input power and output power during the conversion process, and the peak current... This refers to the maximum current that the traction converter can withstand. By adjusting these two parameters, the load on the traction converter can be appropriately adjusted to prevent equipment failure or overheating due to excessive load.
[0103] Secondly, security control disabled flag This refers to whether, under specific operating conditions, the traction converter is allowed to enter emergency shutdown or self-protection mode. When At that time, the system is in normal working mode, while when When this happens, the system's protection functions will be disabled, and the converter will continue to operate under high load until other health management strategies take effect.
[0104] Specifically, the traction converter control unit performs the following operations based on the issued instructions: Adjusting the modulation ratio By increasing the modulation ratio This can improve power conversion efficiency when the traction system is under high load, thereby preventing overload or overheating problems. The modulation ratio can be adjusted by using the PWM (Pulse Width Modulation) control signal of the traction converter to change the duty cycle of the current waveform.
[0105] Adjust peak current The adjustment of peak current directly affects the maximum output power of the converter. When the load on the traction converter is too large, reducing the peak current can effectively reduce current surges, prevent current overload in the system, and protect electrical components from high current damage.
[0106] Disable security controls Under specific operating conditions, such as when the traction converter is in a temporary overload state, disabling the safety control function helps the system continue to operate under extreme conditions, thereby extending the equipment's operating time until it can return to normal operation.
[0107] In this embodiment, experiments were conducted on the traction converter under different operating conditions to execute health management commands. The results show that health management commands can effectively reduce the system failure rate. In one experiment, by reducing the peak current... The system overload rate was reduced by 30%, and the modulation ratio was adjusted. Subsequently, the conversion efficiency of the traction system improved by approximately 15%. This demonstrates the effectiveness of the health management instructions, especially under high load conditions, which effectively protect the equipment and extend the service life of the traction converter.
[0108] In this embodiment, step S10 will be described in detail. In step S10, when the population stability index PSI meets the condition... At that time, the system will activate the incremental learning mechanism and use a fixed sample replay ratio. The sample replay strategy is used to update the model parameters.
[0109] Incremental learning involves adding new data to a model and gradually adjusting existing parameters, enabling the model to adapt to environmental changes and improve predictive performance. Incremental learning is particularly important in traction converter fault prediction because the operating conditions of the traction system are dynamically changing, and fault modes also change over time. Incremental learning allows the model to adjust its predictions in real time after receiving new data, thereby improving prediction accuracy.
[0110] The sample replay strategy uses a fixed ratio By replaying samples selected from historical data, overfitting of the model is avoided while maintaining its ability to learn long-term trends. The replayed samples include historical failure data and normal operation data, which are incorporated into the incremental learning process to ensure the model can continuously update and adapt to new failure modes. A fixed sample replay ratio ensures a balance between old and new data, enabling incremental learning to quickly respond to changes in new data while retaining sufficient historical data to enhance the model's stability and generalization ability.
[0111] The specific steps of incremental learning include: Real-time data collection: The traction converter control system collects operating data in real time, including parameters such as voltage, current, and temperature, and annotates the new data according to the current status of the system.
[0112] Parameter adjustment: When new fault data or operational data is added, the model will adjust the weights and biases based on this data, making the fault prediction model more targeted.
[0113] Model updates: Each update optimizes the weights based on historical data and new measurement data, thereby gradually improving the model's predictive ability.
[0114] In this embodiment, the fault prediction accuracy of the system is significantly improved by introducing incremental learning and sample replay strategies. In one experiment, the fault prediction accuracy improved by 25% after the system adaptively adjusted to new fault data through incremental learning. In addition, the sample replay strategy enables the model to avoid overfitting when facing various different operating conditions, maintain high generalization ability, and reduce the system response time by about 10%.
[0115] It should be noted that the embodiments of the present invention have better implementability and are not intended to limit the present invention in any way. Any person skilled in the art may use the above-disclosed technical content to change or modify it into equivalent effective embodiments. However, any modifications or equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.
Claims
1. A method for fault prediction and health management of intelligent traction converters in rail transit, characterized in that, Includes the following steps: Step S1: Collect the electrical and status quantities of the traction converter control system, including the DC bus voltage. Three-phase current Transient collector-emitter voltage waveform of power semiconductor switch Gate current waveform On-state voltage drop sampling value Gate control signal Device junction temperature estimate and operating condition parameters ,in For train speed, Traction / braking mode; Step S2: Under traction error constraints Below, low-amplitude diagnostic perturbations are injected into the carrier duty cycle. perturbation frequency satisfy carrier frequency Simultaneously record the voltage change caused by the perturbation. With change in current ; Step S3: In the preset frequency band Internal calculation of equivalent series resistance of DC link And extract the gate charge integral. Duration of shutting down the Miller platform and spectral sideband energy ; Step S4: Generate a directed graph of three phases-bridge arm-submodule based on the topology configuration file, construct the channel-submodule topology association matrix M, and then... The M is mapped to a submodule-level feature vector. ; Step S5: Establish a degenerate state-space model, which takes the following form: , , in, State vector; input vector And set the variable to indicate the variable. Threshold for distinguishing change points When the change point is determined by the log-likelihood ratio statistic When the degradation mechanism of the degradation state space model changes, it is determined that the degradation mechanism has changed. Step S6: Perform mean-field variational inference on the degraded state-space model to obtain the failure probability vector. Remaining life estimate and remaining lifetime uncertainty ; Step S7: Perform Dirichlet temperature scaling calibration on the fault occurrence probability vector, with temperature scaling parameters... The calibrated failure probability vector is obtained. Perform an isotropic univariate order-preserving regression calibration on the remaining lifetime estimate to establish a monotonic calibration mapping. The calibrated remaining lifetime estimate is obtained. ; Step S8: Construct the safety barrier function : , , ; One of them, Within the time window The mean, Within the time window standard deviation This refers to the upper limit of the rated peak allowable current of the device. Set for peak current. For safety reasons, For the standby threshold; within the time budget Internally solve the following formula to determine the health management instruction variables. : , ; in , , For weight parameters, The modulation ratio; Step S9: Send the health management command to the traction converter control unit and within the time budget. Internal execution; upper limit of junction temperature Given by the device's rated thermal parameters; Step S10: When the population stability index PSI satisfies At that time, incremental learning is performed and a fixed sample replay ratio is used. The sample replay strategy updates the model parameters; the PSI is calculated using the binning histogram of the reference window and the current window distribution.
2. The method according to claim 1, characterized in that, The diagnostic perturbation uses the jitter mode of pseudo-random binary sequence (PRBS). upper limit of amplitude satisfy and The duration is one carrier cycle and and satisfy .
3. The method according to claim 1, characterized in that, The spectral sideband energy In the set of sideband frequencies The above calculation; the duration of the shut-off Miller platform. The transient collector-emitter voltage waveform of the power semiconductor switch The first-order differential zero-crossing count of the platform segment is obtained; the gate charge integral is obtained. The gate current waveform within one switching cycle Points are earned.
4. The method according to claim 1, characterized in that, The estimated junction temperature of the device is determined by a dual-mode thermal observer; the dual-mode thermal observer includes an equivalent... Type of thermal network parameters , , , The recursive observations, and in Enable with forgetting factor Recursive least squares identification is used to update the equivalent thermal parameters online. , The root mean square of the observed residuals serves as the mode switching criterion.
5. The method according to claim 1, characterized in that, The channel-submodule topological association matrix M is automatically generated based on the directed graph of the phase-arm-submodule; a submodule health status vector is defined. Using group sparsity and The joint regularization degradation coefficient estimation model has the following objective function: , Wherein, the design matrix X is composed of the sub-module level feature vectors Composition, degenerate observation vector Regularization coefficient , Parameter group .
6. The method according to claim 1, characterized in that, The mean field variational inference for the With the The joint posterior distribution is used to solve the coordinate ascent problem, and the failure probability vector is output. The remaining lifetime estimate and its variance The logarithmic probability vector is softmax is the log-probability vector. The class-wise exponential normalization mapping.
7. The method according to claim 1, characterized in that, The Dirichlet temperature scaling calibration acts on the fault occurrence probability vector. The temperature scaling parameter The univariate order-preserving regression calibration is applied to the remaining lifetime estimate. The monotonic calibration mapping is as follows .
8. The method according to claim 1, characterized in that, The health management instructions are optimized using quadratic programming and within the time budget. The constraint set includes the upper limit of the rated peak allowable current of the device. The upper limit of junction temperature and based on the calibrated remaining lifetime estimate The standby threshold ; Output the health management instruction variables .
9. The method according to claim 1, characterized in that, Time synchronization employs a boundary clock method based on a precision time protocol for unified channel calibration, and linear interpolation compensation is applied to the sampling timestamps to ensure that the deviation between the acquisition clock and the control clock meets the required standards. When referenced again, the stated This is the synchronization accuracy threshold.
10. The method according to claim 1, characterized in that, When the PSI exceeds the trigger threshold And the change point discriminant log-likelihood ratio statistic Exceeding the threshold When the H-infinity robust state estimator is enabled, the robust term coefficient... Its optimized formula is Output the probability vector of the backup failure occurrence. Compared with the estimated remaining spare life The primary inference and the backup inference are selected based on an uncertainty threshold; wherein the uncertainty radius is... .