Method and system for health management of power module

CN122652376APending Publication Date: 2026-08-28INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611135813.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0003]本申请提供了一种电源模块的健康管理方法及系统,以至少解决相关技术中电源寿命预测的预测结果难以准确反映电源模块在实际设备工作条件下的真实剩余使用寿命的问题

Benefits of technology

[0008] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described power module health management methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122652376A_ABST
    Figure CN122652376A_ABST
Patent Text Reader

Abstract

The application discloses a health management method and system of a power module, relates to the technical field of electricity, and can evaluate the load frequent fluctuation, environment temperature dynamic change and other working condition characteristics faced by the device power supply in actual operation by synchronously collecting thermal, electrical and dynamic response data, calculate the health characteristic index representing the aging characteristics of the power module based on the multi-dimensional operation parameters, and input the health index obtained by a light-weight model, so as to avoid the environmental noise and working condition interference introduced by directly using the original data. Further, the attenuation trajectory is predicted according to the health index sequence, compared with the preset failure threshold to determine the target failure time and calculate the remaining service life, so that the prediction result reflects the dynamic attenuation process of the health index, and the problem that the single-point evaluation is not sensitive to the change of the attenuation rate is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electrical technology, and in particular to a health management method and system for a power module. Background Technology

[0002] During operation, the power modules of electronic devices experience performance degradation due to long-term exposure to electrothermal stress. When this degradation accumulates to a certain extent, it can lead to power module failure. Therefore, it is necessary to accurately predict the remaining lifespan of the power supply to avoid equipment downtime caused by sudden power failures. However, power supply lifespan prediction schemes in related technologies usually only analyze the electrical parameters of the power module itself and fail to make predictions based on the actual operating conditions of the power supply. This makes it difficult for the prediction results to accurately reflect the true remaining lifespan of the power module under actual equipment operating conditions. Summary of the Invention

[0003] This application provides a health management method and system for power modules, which at least solves the problem in the related art that the prediction results of power life prediction are difficult to accurately reflect the actual remaining service life of the power module under actual equipment working conditions.

[0004] This application provides a health management method for a power module, including: At multiple designated monitoring points corresponding to the power module of the target device, real-time operating parameters reflecting the aging status of the designated monitoring points are collected synchronously. The real-time operating parameters include at least the thermal data, electrical data and dynamic response data of the power module. Health characteristic indicators that characterize the aging of the power module are calculated based on real-time operating parameters, and the health characteristic indicators are input into a pre-trained lightweight health index evaluation model to obtain the current health index of the power module at the current moment. The decay trajectory of the current health index is predicted based on the sequence data corresponding to the current health index, and the predicted health trajectory is obtained. The sequence data is a health index sequence that is constructed based on the current health index and the historical health index, and the health index changes over time. The predicted trajectory is compared with the preset failure index threshold to determine the target failure time, and the time difference between the target failure time and the current time is determined as the remaining service life of the power module.

[0005] This application also provides a health management system for a power module, including: The parameter acquisition module is used to synchronously acquire real-time operating parameters reflecting the aging status of multiple designated monitoring points corresponding to the power module of the target device. The real-time operating parameters include at least the thermal data, electrical data and dynamic response data of the power module. The feature calculation module is used to calculate health feature indicators that characterize the aging characteristics of the power supply module based on real-time operating parameters. An embedded processing module is used to input health characteristic indicators into a pre-trained lightweight health index evaluation model to obtain the current health index of the power module at the current moment. The lifespan prediction module is used to predict the decline trajectory of the current health index based on the sequence data corresponding to the current health index, and obtain the predicted health trajectory. The sequence data is a health index sequence that is constructed based on the current health index and the historical health index, and the health index changes over time. The life prediction module is also used to compare the predicted trajectory with the preset failure index threshold to determine the target failure time, and to determine the time difference between the target failure time and the current time as the remaining life of the power module. The data communication interface is used to report the current health index and remaining service life to the controller or remote monitoring platform of the target device.

[0006] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described power module health management methods when executing the computer program.

[0007] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described power module health management methods.

[0008] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described power module health management methods.

[0009] The power module health management method and system of this application, by simultaneously collecting thermal, electrical, and dynamic response data, can assess the power supply based on the operating conditions faced by the device power supply in actual operation, such as frequent load fluctuations and dynamic changes in ambient temperature. Simultaneously, it calculates health characteristic indicators representing the aging characteristics of the power module based on multi-dimensional operating parameters and inputs them into a lightweight model to obtain a health index. This avoids the environmental noise and operating condition interference introduced by directly using raw data. Furthermore, by predicting the attenuation trajectory based on the health index sequence and comparing it with a preset failure threshold to determine the target failure time and estimate the remaining service life, the prediction results reflect the dynamic attenuation process of the health index, solving the problem of single-point assessment being insensitive to changes in the attenuation rate. Therefore, this application can directly achieve high-precision online prediction of the remaining service life within the power module, and the prediction results can truly reflect the state and remaining service life of the power module under actual equipment operating conditions. Attached Figure Description

[0010] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating a health management method for a power module provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a health management system for a power module provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0013] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0014] This application provides a health management method for power modules, used for real-time monitoring and lifespan prediction of power module degradation status, thereby providing a basis for predictive maintenance decisions for electronic devices. As electronic devices operate for longer periods, their internal power modules are subjected to the combined effects of various factors such as electrothermal stress, ambient temperature fluctuations, and frequent load switching, leading to slow and irreversible performance degradation of internal power devices, magnetic components, and energy storage components. If the health status of the power module can be accurately detected in the early stages of degradation, and its remaining lifespan can be predicted before complete failure, it can effectively prevent equipment downtime, data loss, and high maintenance costs caused by sudden power failures.

[0015] The method provided in this application can run in the embedded controller built into the power module, or it can run in the core controller of the target device. This application does not specifically limit it in this way. For ease of description, the embodiments of this application mainly use the example of the method running in the microcontroller unit (MCU) built into the power module for illustration.

[0016] A power module is a power conversion unit used to provide stable power to the core loads of a target device, such as the Central Processing Unit (CPU), Graphics Processing Unit (GPU), memory, and storage devices. A power module may include a power factor correction (PFC) stage, a DC-to-DC converter stage, an output filtering stage, and corresponding control and protection circuits.

[0017] The target device refers to the device that requires power health management. This target device includes, but is not limited to, servers, switches, personal computers, new energy vehicles, etc. Specifically, this application does not limit the target device.

[0018] To meet the high power density and high reliability requirements of target equipment, power modules often adopt a multi-phase parallel topology. Each phase consists of power switching devices, power inductors and corresponding drive circuits. The multiple phases operate alternately with a certain phase difference to jointly deliver current to the load.

[0019] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] Figure 1 This document provides a flowchart illustrating a health management method for a power module, as provided in an embodiment of this application. The method is described in detail below, taking into account the execution flow of the health management method for the power module.

[0021] Step 101: Simultaneously collect real-time operating parameters reflecting the aging status of the specified monitoring points at multiple designated monitoring points corresponding to the power module of the target device. The real-time operating parameters include at least the thermal data, electrical data and dynamic response data of the power module.

[0022] In the embodiments of this application, a designated monitoring point refers to a physical location or electrical node on the power module that can reflect the performance degradation process of internal key components. The selection of the designated monitoring point is determined based on the failure physics analysis of the power module, that is, by analyzing the failure modes and failure mechanisms of the power module under various stress conditions, the components most prone to aging and performance drift and their measurable characterization parameters are identified.

[0023] In actual implementation, designated monitoring points may include, but are not limited to: the drain-source terminals (or collector-emitter terminals) of power switching devices, the drive pins of power switching devices and the common terminal of the power circuit, the two ends of the output filter capacitor, the winding taps of the power inductor and the magnetic core grounding terminal, the positive and negative terminals of the input bus of the power module, the positive and negative terminals of the output bus, and temperature measurement points at key heat dissipation locations inside the power module.

[0024] Synchronous data acquisition refers to the simultaneous sampling of different physical quantities at multiple specified monitoring points at the same sampling time or within a very short time window. Since the internal electrical state of a power module changes rapidly and dynamically with load variations during operation, synchronous data acquisition ensures that the instantaneous values ​​of each physical quantity are obtained under the same load conditions. This allows for accurate calculation of the true correspondence between parameters and avoids calculation errors introduced by inconsistent sampling times.

[0025] Thermal data refers to physical quantities used to characterize the temperature state of key components or key locations inside a power module. Thermal data can reflect the thermal stress level and heat dissipation status of the power module.

[0026] For example, thermal data may include, but is not limited to: junction temperature of power switching devices, case temperature of power switching devices, temperature of power inductor windings, temperature of power inductor core, surface temperature of output filter capacitor, internal air temperature of power module, and ambient temperature at power module air inlet.

[0027] Electrical data refers to physical quantities used to characterize the electrical characteristics of key nodes in a power module, such as voltage, current, and resistance. The evolution of electrical parameters reflects the degree of degradation of components within the power module.

[0028] For example, electrical data may include, but is not limited to: the drain-source voltage and on-current of the power switching device in the fully on state, from which the on-resistance can be calculated; the leakage current of the power switching device in the off state; the DC voltage across the output filter capacitor; the ripple voltage and ripple current of the output filter capacitor; the average current and ripple current flowing through the power inductor; the input DC voltage and input current of the power module; the output DC voltage and output current of the power module; and the equivalent series resistance value of the output filter capacitor.

[0029] Dynamic response data refers to physical quantities used to characterize the output voltage recovery capability and response characteristics of a power module when the load conditions undergo a step change. As the internal components of the power module age, the dynamic response performance of the power module will gradually deteriorate. Therefore, dynamic response data can reflect the health degradation status of the power module from a time domain perspective.

[0030] For example, dynamic response data can be obtained by capturing the instantaneous waveform of the output voltage when the load current changes stepwise during the operation of the power module using an analog comparator or a high-speed analog-to-digital converter (ADC) inside the embedded controller. Specifically, this may include, but is not limited to: the drop in output voltage (undershoot) when the load current steps from light load to heavy load, the overshoot of output voltage when the load current steps from heavy load to light load, the time required for the output voltage to recover to the specified error range of the steady-state value from the occurrence of the disturbance (recovery time), and the deviation of the steady-state value of the output voltage before and after the load step from the rated value (steady-state error).

[0031] Step 102: Calculate the health characteristic index that characterizes the aging characteristics of the power module based on the real-time operating parameters, and input the health characteristic index into the pre-trained lightweight health index evaluation model to obtain the current health index of the power module at the current moment.

[0032] In the embodiments of this application, the health characteristic index refers to the comprehensive quantitative value that can quantitatively reflect the overall or local aging degree of the power module from different perspectives after fusing and calculating various real-time operating parameters.

[0033] Individual raw data acquisition parameters (such as on-state voltage drop and absolute value of ripple voltage) are often affected by external operating conditions such as load current, input voltage, and ambient temperature. Changes in their absolute values ​​cannot directly and reliably reflect the degree of aging. Therefore, it is necessary to normalize, compare, and perform comprehensive calculations on the raw parameters to eliminate interference caused by fluctuations in external operating conditions and extract the feature quantities that are truly strongly correlated with internal physical degradation.

[0034] For example, health characteristics may include, but are not limited to: power stage efficiency degradation rate, output voltage ripple increment, dynamic load response degradation, and phase-to-phase current imbalance.

[0035] A lightweight health index assessment model refers to a computational model pre-trained using machine learning or deep learning methods to map health characteristic indicators to health indices. To ensure efficient operation on the embedded controller of the power module, the model adopts a lightweight design, meaning its network structure is relatively simple, the number of parameters is small, and the computational complexity is low, enabling it to perform real-time inference under the limited computing power and storage resources of the embedded controller.

[0036] For example, the lightweight health index assessment model can employ a fully connected neural network structure comprising an input layer, several hidden layers, and an output layer. The number of neurons in the hidden layers can be configured according to the memory capacity and accuracy requirements of the embedded controller. Alternatively, the lightweight health index assessment model can employ a lightweight neural network structure integrating depthwise separable convolutions and gated recurrent units. The depthwise separable convolutions extract local variation patterns of health feature indicators at different time scales, while the gated recurrent units capture the temporal dependencies between health feature indicators. However, this application does not impose specific limitations on this approach; any lightweight model structure capable of mapping health feature indicators to a health index and meeting embedded deployment requirements can be applied to this embodiment.

[0037] The health index is a scalar value used to quantitatively represent the remaining health level of a power module relative to its initial state at a given moment. The numerical range of the health index can be determined according to the design of the model's output layer. For example, the health index can be set to a continuous value between 0 and 1, where a value of 1 indicates that the power module is in a completely healthy state, and a value of 0 indicates that the power module has completely failed. As the power module's operating time increases and its internal components age, the health index shows a gradual decay trend, and its decay rate reflects the current degradation rate of the power module. In another implementation, the health index can also be represented using a percentage system (e.g., 0 to 100) or other normalized numerical ranges; this application does not limit this.

[0038] Step 103: Predict the decay trajectory of the current health index based on the sequence data corresponding to the current health index to obtain the predicted health trajectory. The sequence data is a health index sequence constructed based on the current health index and the historical health index, showing the change of the health index over time.

[0039] In the embodiments of this application, sequence data refers to a health index sequence arranged in chronological order by sorting the current health index and historical health index according to their respective corresponding times. Historical health index refers to the health index values ​​obtained from multiple historical sampling times prior to the current time.

[0040] It should be noted that the distribution of health index values ​​in the sequence data along the time axis is not uniform. The specific time interval depends on the execution frequency of steps 101 and 102, which can be adaptively adjusted according to the drastic changes in the target device's load and the computational load of the embedded controller. In scenarios where the target device's load changes rapidly, a higher sampling frequency can be used to obtain a denser health index sequence to improve prediction accuracy; in scenarios where the target device's load is relatively stable, the sampling frequency can be appropriately reduced to save computational resources.

[0041] The decay trajectory refers to the predicted trend curve of the health index as it changes over time. Since the health index of a power module does not decrease linearly and monotonically throughout its entire life cycle, but may exhibit a nonlinear pattern of slow degradation in the early stage, accelerated degradation in the middle stage, and rapid degradation in the final stage, it is necessary to dynamically estimate the future decay path of the health index based on the recent degradation trend and rate of change implied in the sequence data.

[0042] A predicted health trajectory refers to the set of predicted values ​​of the health index at various points in time over a future period, obtained by extrapolating the decay trajectory using a prediction algorithm. Predicted health trajectories can be obtained using various methods such as curve fitting, extrapolation, or state estimation. For example, multinomial fitting can be used to fit the sequence data and then extrapolate it to the future; Kalman filtering can also be used to recursively estimate the state of the health index and predict its future evolution trend.

[0043] Step 104: Compare the predicted trajectory with the preset failure index threshold to determine the target failure time, and determine the time difference between the target failure time and the current time as the remaining service life of the power module.

[0044] In the embodiments of this application, the preset failure index threshold refers to a pre-set health index threshold used to determine whether the power module has lost its normal operating capability. This preset failure index threshold can be determined comprehensively based on the power module's technical specifications, the target device's power supply quality requirements, and industry standards.

[0045] For example, when the health index of the power module is 80% of that of a brand new module, it indicates that the power module is still in a healthy range and its performance indicators still meet the specifications. When the health index drops to 50% of that of a brand new module, it can be set as a warning zone, indicating that the power module has shown significant aging and a maintenance plan is recommended. When the health index drops to 20% of that of a brand new module, it can be set as a failure threshold, indicating that the power module is on the verge of failure or its output performance no longer meets the power supply requirements of the target device. The specific value of the preset failure index threshold can be calibrated based on the actual test data of different models of power modules, and this application does not impose specific limitations on it.

[0046] The comparison process involves sequentially comparing the predicted health index values ​​for each future moment on the predicted health trajectory with a preset failure index threshold to determine the moment when the predicted health trajectory intersects with the failure threshold line. The target failure moment is the moment when the predicted health trajectory first reaches the preset failure index threshold. This target failure moment represents the earliest moment when the power module is expected to reach a failure state after prediction based on sequence data.

[0047] Remaining Useful Life (RUL) refers to the length of time elapsed from the current moment until the target failure moment, indicating how long the power module is expected to continue to operate normally in its current degraded state.

[0048] The remaining service life can be expressed in hours, days or other time units. This remaining service life can be reported to the controller of the target device or the remote operation and maintenance monitoring platform, so that operation and maintenance personnel can reasonably arrange maintenance windows based on the predicted value of the remaining service life before the power module actually fails, prepare replacement parts in advance and perform replacement operations, thereby minimizing the risk of sudden downtime caused by power module aging failure.

[0049] This application achieves the synchronous acquisition of multi-dimensional operating parameters, extraction of health characteristic indicators, and evaluation of health indices, culminating in the prediction of decay trajectories and the estimation of remaining lifetime based on health index sequences. By synchronously acquiring real-time operating parameters from multiple dimensions, including thermal, electrical, and dynamic response data, the health characteristic indicators can characterize the internal degradation state of the power module from multiple physical levels—thermal, electrical, and dynamic—improving the comprehensiveness and accuracy of health status assessment. Simultaneously, through a lightweight health index evaluation model, the evaluation process can be completed in real-time within the embedded controller of the power module, without relying on external devices or cloud computing resources. This reduces communication latency and dependence on network bandwidth, improving the real-time performance and reliability of lifetime prediction. Furthermore, by using time-varying sequence data of the health index for decay trajectory prediction, the prediction process can utilize the evolutionary trend information of the health index over time, improving the stability and anti-interference capability of the prediction results.

[0050] In one possible implementation of this application embodiment, the acquisition of real-time operating parameters can be achieved in the following ways, but not limited to: When the power switching device is in a fully on state, the voltage and current across the power switching device are simultaneously sampled to obtain the on-state voltage drop of the power switching device. The multiple designated monitoring points include at least the power switching device, the output filter capacitor, and the power inductor. The junction temperature of the power switching device is obtained through a temperature sensor integrated into the power switching device package, or calculated using the on-state voltage drop and a thermal resistance model incorporating known temperature characteristics. A micro-amplitude AC disturbance signal of a preset frequency is injected into the output terminal of the power module, and the equivalent series resistance of the output filter capacitor is calculated based on the phase difference and amplitude ratio of the output voltage ripple and output current ripple caused by the micro-amplitude AC disturbance signal. The electrical data includes at least the on-state voltage drop and junction temperature of the power switching device, the equivalent series resistance of the output filter capacitor, and the output voltage ripple and output current ripple.

[0051] In the embodiments of this application, the fully conducting state of the power switching device means that after receiving a valid drive signal, the internal conductive channel of the power switching device is fully opened, the drain-source voltage is reduced to a minimum value and maintains a stable operating state.

[0052] Taking a metal-oxide-semiconductor field-effect transistor (MOSFET) as an example, when its gate drive voltage exceeds the threshold voltage and reaches a sufficiently enhanced voltage, the device enters the linear region (also known as the variable resistance region). At this point, the ratio of drain-source voltage to drain current is approximately equal to the on-resistance. Since the power switching device alternates between the on and off states during the switching cycle, only when sampling is performed during the duration of its fully on state can the measured voltage truly reflect the on-state resistance characteristics of the device—the on-state voltage drop.

[0053] Since the on-state voltage drop is equal to the product of the current and the on-state resistance, the on-state resistance can be indirectly calculated by Ohm's law after sampling the on-state voltage drop and the on-state current. The change in the on-state resistance directly reflects the degree of degradation of conductivity caused by the accumulation of lattice damage and interface state traps inside the device.

[0054] In this embodiment, the junction temperature of the power switching device can be obtained using one of two methods or a combination thereof.

[0055] The first method involves directly obtaining the junction temperature using a temperature sensor integrated into the power switching device package. The power switching device integrates a thermistor, i.e., a temperature sensor, for temperature detection within the same package at the factory. This thermistor may include, but is not limited to, a thermistor diode, a thermistor resistor, or a temperature sensing circuit. In one embodiment, the embedded controller can directly obtain the measured junction temperature value by reading the output signal of this integrated temperature sensor, performing analog-to-digital conversion, or reading it through a communication interface.

[0056] The second method involves calculating the junction temperature of the power switching device using the on-state voltage drop and a thermal resistance model that incorporates known temperature characteristics. The on-state voltage drop of a semiconductor power device exhibits a certain functional relationship with the junction temperature under constant on-state current. For example, for a MOSFET device, its on-resistance increases with increasing junction temperature (positive temperature coefficient characteristic), which can be obtained from the temperature-on-resistance characteristic curve provided in the device datasheet.

[0057] Based on this physical characteristic, the junction temperature can be calculated using a pre-established thermal resistance model. The thermal resistance model is a mathematical expression used to describe the relationship between the temperature difference and heat dissipation power along the heat conduction path inside a power switching device. Its basic form is: Junction temperature = Case temperature + Device heat dissipation power × Junction-to-case thermal resistance.

[0058] Based on the on-state voltage drop and on-state current obtained through synchronous sampling, the conduction loss of the device can be calculated. Combined with the case temperature acquired by a temperature sensor and the junction-to-case thermal resistance parameters provided in the device datasheet, the junction temperature can be calculated. These two methods of obtaining the junction temperature can complement each other. In the event of a faulty or missing integrated temperature sensor, the thermal resistance model calculation method can be used as a substitute. If the thermal resistance model parameters are inaccurate, direct measurement by the temperature sensor can be used for calibration.

[0059] In this embodiment, the method for acquiring the equivalent series resistance of the output filter capacitor specifically adopts a combination of micro-amplitude AC disturbance injection and response detection. The equivalent series resistance is a parameter that measures the degree of performance degradation of the output filter capacitor. It represents the equivalent resistance component of the capacitor's internal electrodes, electrolyte, and leads. When AC current flows through the capacitor, this equivalent series resistance will generate heat dissipation and cause an additional voltage drop, thereby increasing the output voltage ripple.

[0060] A small-amplitude AC disturbance signal refers to an AC sinusoidal signal with a small amplitude that will not affect the normal operating output voltage of the power module. The frequency of this small-amplitude AC disturbance signal is preset to a frequency value that allows the output filter capacitor to exhibit measurable capacitive reactance and resistive characteristics. For example, this preset frequency can be set to a frequency point much lower than the switching frequency of the power switching device.

[0061] A small-amplitude AC disturbance signal can be generated by a disturbance signal generating circuit connected in series at the output of the power module. Output voltage ripple refers to the AC voltage component superimposed on the output DC voltage, and output current ripple refers to the AC current component superimposed on the output DC current.

[0062] When a small AC disturbance signal is injected into the output terminal, the disturbance signal forms an AC loop through the output filter capacitor and the load resistor, generating corresponding AC voltage and AC current responses across the capacitor. By detecting the phase difference and amplitude ratio between the AC voltage and current responses, the impedance modulus and impedance angle of the capacitor at the preset frequency can be calculated, and then the specific value of the equivalent series resistance can be derived based on the series equivalent circuit model of the capacitor.

[0063] The electrical data includes at least the on-state voltage drop and junction temperature of the power switching devices, the equivalent series resistance of the output filter capacitor, and the output voltage ripple and output current ripple. It should be noted that the output voltage ripple and output current ripple can be detected as response signals during the injection of a small-amplitude AC disturbance signal, or they can be obtained by AC coupling sampling at the output terminal during normal operation of the power module under load; this application does not limit the specific method used.

[0064] This application achieves precise acquisition of electrical data for key components of the power module by simultaneously acquiring the on-state voltage drop and on-state current of the power switching device while it is fully on, obtaining the junction temperature through a temperature sensor or thermal resistance model, and measuring the equivalent series resistance of the output filter capacitor using a micro-amplitude AC perturbation injection method. This allows key electrical parameters such as on-state voltage drop, junction temperature, and equivalent series resistance to be acquired online without affecting the normal power supply of the power module, eliminating the need to remove the power module from the equipment for offline testing and ensuring the continuity of equipment operations.

[0065] In one possible implementation of this application, the thermal data includes at least the winding temperature and core temperature of the power inductor, and the dynamic response data includes at least the ambient temperature and load current change rate of the power module.

[0066] In the embodiments of this application, the power inductor is the core magnetic component in the power module responsible for the conversion between magnetic energy and electrical energy. Its performance stability directly affects the conversion efficiency and output characteristics of the power module. Winding temperature refers to the temperature of the conductors winding the coil in the power inductor, specifically the average temperature of the winding or the hot spot temperature of the winding. Core temperature refers to the temperature of the core of the power inductor. The core is typically made of soft magnetic materials such as ferrite, metal powder core, or amorphous nanocrystals, and its magnetic parameters, such as permeability, saturation flux density, and coercivity, are significantly affected by temperature.

[0067] Ambient temperature refers to the air temperature at the inlet of the power module's operating environment, i.e., the temperature of the cooling airflow before it enters the power module. Load current change rate refers to the rate at which the power module's output current changes over time; specifically, it can be understood as the derivative or difference of the load current with respect to time. This parameter is used to quantitatively reflect the drastic degree of load change in the target device.

[0068] In one possible implementation of this application embodiment, when calculating the health characteristic indicators characterizing the aging characteristics of the power module, the following methods can be used, but are not limited to: calculating the efficiency degradation rate of the power module based on the input voltage, input current, output voltage, and output current in the real-time operating parameters; calculating the output voltage ripple increment of the power module based on the equivalent series resistance and current ripple of the output filter capacitor in the real-time operating parameters; calculating the dynamic load response degradation degree of the power module based on the on-state voltage drop and junction temperature of the power switching devices, and the winding temperature and core temperature of the power inductor in the real-time operating parameters; and calculating the phase-to-phase current imbalance degree of the power module based on the current data in the real-time operating parameters.

[0069] In the embodiments of this application, the efficiency degradation rate is an indicator that measures the degree of attenuation of the power conversion capability of the power module. As the operating time of the power module accumulates, the on-resistance of its internal power switching devices increases, leading to increased switching and conduction losses. The winding resistance and core loss of the power inductor increase, resulting in increased magnetic component losses. The equivalent series resistance of the output filter capacitor increases, leading to increased ripple losses. The cumulative effect of various losses ultimately manifests as a decrease in the overall power conversion efficiency of the power module.

[0070] Output voltage ripple increment is an indicator of the degree of degradation in the filtering performance of the output filter capacitor. The output filter capacitor is an energy storage component at the output terminal of the power module used to smooth the output voltage. Its equivalent series resistance and capacitance value together determine the filtering effect at the output terminal. When the equivalent series resistance gradually increases during long-term operation due to electrolyte drying or dielectric material aging, the high-frequency ripple current flowing through the capacitor generates a larger voltage drop across the equivalent series resistance, thereby increasing the voltage ripple amplitude at the output terminal.

[0071] Specifically, the output voltage ripple increment = the effective value of the current ripple flowing through the output filter capacitor × the increment of the equivalent series resistance relative to its brand-new state. The calculated output voltage ripple increment reflects the ripple voltage amplitude additionally superimposed on the output voltage due to capacitor aging. A larger output voltage ripple increment indicates more severe degradation of the output filter capacitor's filtering performance and poorer output voltage quality. It is worth noting that since the magnitude of the current ripple is directly related to the load current, when calculating the output voltage ripple increment, the actual measured current ripple can be converted to the equivalent current ripple under rated load according to the current load rate, and then multiplied by the equivalent series resistance increment to eliminate the influence of load changes on the ripple increment calculation result.

[0072] Dynamic load response degradation is a comprehensive indicator that measures the degree of degradation in the output regulation capability and response speed of a power module when the load undergoes a step change. When the service load of the target device changes abruptly, the output current of the power module increases or decreases instantaneously, and the output voltage will have a certain instantaneous deviation due to the impedance characteristics of the output filter capacitor and the regulation capability of the control loop. As the internal components of the power module age, its dynamic response performance will gradually deteriorate. Specifically, an increase in the on-state voltage drop and junction temperature of the power switching devices leads to a slower switching speed and a decrease in drive capability, making it impossible for the power switching devices to quickly respond to the adjustment of the control signal during load transients; an increase in the winding temperature and core temperature of the power inductor leads to inductance drift and changes in saturation characteristics, reducing the inductor's ability to store and release magnetic energy during transients. Together, these factors result in an increase in the overshoot and undershoot of the output voltage of the power module during load step changes, as well as a prolonged recovery time.

[0073] Current imbalance between phases is a quantitative indicator measuring the uniformity of current distribution among phases in a power module employing a multiphase parallel topology. For a target device power module with a multiphase parallel topology, each phase power stage operates alternately with a certain phase difference, jointly providing current to the load. Ideally, the current in each phase should be evenly distributed, meaning the current amplitude in each phase is essentially equal. However, as the power module operates over time, the degradation rates of components such as power switches and power inductors in each phase power stage may differ. For example, if the on-resistance of a power switch in one phase increases rapidly, the current carrying capacity of that phase decreases, requiring other phases to carry a larger share of current. Similarly, a large inductance drift in a power inductor in one phase may cause its current ripple to be inconsistent with other phases. This imbalance further exacerbates the heating and aging of the phase with the higher current carrying capacity, creating a positive feedback acceleration effect, which may ultimately lead to the failure of one phase and trigger a failure of the entire power module.

[0074] This application achieves the assessment of the aging status of power modules from the dimensions of efficiency, filtering performance, dynamic response, and current balance through the calculation of health characteristic indicators. Each health characteristic indicator starts from different physical mechanisms and emphasizes the degradation sensitivity of different components. Combined, they form a highly complementary health characteristic vector, providing complete information for the accurate assessment of the overall health index of the power module by the lightweight health index assessment model.

[0075] In one possible implementation of this application embodiment, when calculating the efficiency degradation rate of the power module, it can be implemented in the following ways, but not limited to: calculating the current efficiency of the power module based on the input voltage, input current, output voltage, and output current in the real-time operating parameters; comparing the current efficiency with the reference efficiency of the power module to obtain the percentage decrease of the current efficiency relative to the reference efficiency, and determining the percentage decrease as the efficiency degradation rate, wherein the reference efficiency is measured by the power module within the historical calibration period or obtained by preset.

[0076] In the embodiments of this application, the current efficiency of the power module is a performance indicator that measures the current power conversion capability. It is the ratio of output power to input power and reflects the degree of energy loss in the process of converting input electrical energy into output electrical energy.

[0077] In the specific calculation, the instantaneous values ​​of the input voltage and input current are acquired synchronously, and multiplied together to obtain the input power. Simultaneously, the instantaneous values ​​of the output voltage and output current are acquired synchronously, and multiplied together to obtain the output power. After obtaining the input and output power, the output power is divided by the input power to obtain the current efficiency of the power module under the current operating conditions. It should be noted that the efficiency of the power module varies under different load rates, generally reaching its highest efficiency within the 50% to 80% rated load range, while efficiency decreases at light and heavy loads. Therefore, when calculating the current efficiency, it is necessary to record the load rate corresponding to the current moment (i.e., the ratio of output current to rated output current) so that subsequent comparisons can be performed under the same load rate benchmark.

[0078] Reference efficiency refers to the baseline efficiency value of a power module when it is in good health, serving as a zero point for judging the degree of efficiency degradation. Reference efficiency is measured by the power module within a historical calibration period or obtained through a preset method. Specifically, the historical calibration period refers to the calibration time point prior to the current moment, at which the power module is considered to be in a healthy or known performance state. The efficiency value of the power module is measured under standard test conditions and recorded as the reference efficiency.

[0079] For example, the historical calibration cycle can be an efficiency test performed during the factory testing phase when the power module leaves the factory, and the measured efficiency value is written into the non-volatile memory of the embedded controller as the factory reference efficiency. Alternatively, the historical calibration cycle can also be an efficiency test automatically performed during the initial power-on self-test phase after the power module is put into operation in the target device, and the measured efficiency value is stored as the initial reference efficiency. Or, the historical calibration cycle can also be an efficiency calibration test that is automatically performed periodically during idle periods of the power module during operation, and the efficiency value obtained from each calibration test can be updated to a new reference efficiency to eliminate the cumulative error caused by slow drift (such as temperature sensor bias drift, changes in sampling resistor value, etc.) other than natural aging of the power module in calculating the efficiency degradation.

[0080] In addition, the baseline efficiency can also be obtained through a preset method. The preset method refers to the efficiency baseline value determined in advance by theoretical calculation or batch test statistics based on the typical efficiency characteristic curve of the same model of power module under standard test conditions. This preset baseline efficiency is fixed in the storage medium of the embedded controller before the power module leaves the factory.

[0081] When comparing the current efficiency with the baseline efficiency, the load rate corresponding to the current efficiency is first determined. Then, the baseline efficiency value corresponding to the same load rate is obtained from the baseline efficiency data. The percentage decrease in current efficiency relative to this baseline efficiency is calculated. The specific formula for calculating the percentage decrease is: Efficiency degradation rate = (Baseline efficiency value - Current efficiency value) / Baseline efficiency value × 100%. When the power module's efficiency has not degraded, the current efficiency equals the baseline efficiency, and the efficiency degradation rate is 0%; when the power module's efficiency has completely degraded, the efficiency degradation rate approaches 100%.

[0082] In practical applications, different types of power modules have different maximum allowable efficiency degradation rates, which are usually related to the power module's technical specifications and the target device's minimum power efficiency requirements. For example, for some high-efficiency target device power modules, when their efficiency degradation rate exceeds 3%, it can be determined that their health status has significantly declined, and maintenance and inspection should be arranged.

[0083] This application achieves a quantitative characterization of the degree of efficiency degradation by defining the percentage decrease between the current efficiency and the baseline efficiency as the efficiency degradation rate, which helps the model to more accurately assess the current health index of the power module. Furthermore, since the calculation of the efficiency degradation rate only depends on the input voltage, input current, output voltage, and output current, no additional hardware sensors are required. It can be implemented using the existing voltage and current sampling circuits of the power module, offering advantages such as simple calculation and no increase in hardware cost.

[0084] In one possible implementation of this application embodiment, the calculation of the dynamic load response degradation degree of the power module can be achieved in the following ways, but not limited to: performing a preset step load change test on the power module to obtain the overshoot amplitude, recovery time, and steady-state error of the output voltage; performing temperature compensation correction on the overshoot amplitude, recovery time, and steady-state error based on the on-state voltage drop and junction temperature of the power switching device and the winding temperature and core temperature of the power inductor to obtain the corrected response parameters; performing weighted calculation on the overshoot amplitude, recovery time, and steady-state error in the corrected response parameters based on preset weights to obtain the target response curve; and comparing the target response curve with the preset standard response curve to obtain the dynamic load response degradation degree.

[0085] In the embodiments of this application, the preset step load change test refers to a dynamic test program with the load change amplitude set in advance.

[0086] Overshoot refers to the maximum positive deviation of the output voltage from its target steady-state value after a load step change. Recovery time refers to the time elapsed from the moment the load step change occurs until the output voltage recovers to its steady-state value and remains within the specified error range of the steady-state value. Steady-state error refers to the deviation between the steady-state value reached when the output voltage finally stabilizes after the load step change is completed and its target rated value.

[0087] After obtaining the raw measurements of overshoot amplitude, recovery time, and steady-state error, a temperature compensation correction step is further performed. Because the dynamic response performance of the power module is significantly affected by temperature, parameters such as the switching speed, on-resistance, and drive threshold voltage of power switching devices change under different temperature conditions. The permeability and winding resistance of power inductors also drift with temperature changes. All these factors have a cumulative impact on the dynamic response test results. If the effects of temperature changes and aging degradation are not distinguished when calculating the degree of dynamic load response degradation, temporary performance changes caused by temperature will be misjudged as permanent aging degradation, leading to inaccurate health status assessments.

[0088] Temperature compensation correction refers to converting the measured dynamic response parameters to equivalent values ​​under standard temperature conditions based on currently acquired temperature data. Specifically, the implementation involves: obtaining temperature-dependent characteristic curves of the power module's dynamic response parameters under different temperature conditions through pre-testing experiments. These curves record the corresponding changes in overshoot amplitude, recovery time, and steady-state error for each unit change in temperature. During actual compensation, one or more of the currently acquired junction temperature, case temperature, winding temperature, and core temperature are used as input. The corresponding temperature compensation coefficient is obtained by searching the temperature-dependent characteristic curve. Then, the measured overshoot amplitude, recovery time, and steady-state error are multiplied by or added to their respective temperature compensation coefficients to obtain the corrected response parameters after temperature normalization.

[0089] The preset weights are pre-defined weighting coefficients used to reflect the relative importance of overshoot amplitude, recovery time, and steady-state error in evaluating dynamic response performance. In practical applications, different target device load types have varying sensitivities to dynamic response parameters. For example, for CPU loads sensitive to voltage sags, overshoot amplitude (especially undershoot amplitude) has a more significant impact on system stability, thus overshoot amplitude can be assigned a larger weight; for business scenarios requiring rapid recovery, recovery time can be assigned a larger weight; and for loads requiring high-precision power supply, steady-state error can be assigned a larger weight. The specific values ​​of the preset weights can be flexibly configured according to different application scenarios and the power quality requirements of the target device; this embodiment does not impose specific limitations on this.

[0090] The preset standard response curve refers to the dynamic response benchmark data obtained when the power module is in a brand new, healthy state and undergoes a step load change test with the same parameters under standard temperature conditions. The preset standard response curve can be obtained through laboratory calibration before the power module leaves the factory.

[0091] The comparison between the target response curve and the preset standard response curve can be achieved in various ways. For example, the difference between the overall score of the target response curve and the overall baseline value of the preset standard response curve can be quantified as the dynamic load response degradation degree; a larger difference indicates a more severe degradation in dynamic response performance. More specifically, the percentage deviation of the overall score of the target response curve from the overall baseline value of the preset standard response curve can be calculated; this percentage value is the dynamic load response degradation degree. A higher dynamic load response degradation degree indicates a more significant decrease in the power module's ability to maintain stable output voltage during transient load changes compared to its initial state.

[0092] The application utilizes compensation and correction based on the junction temperature of power switching devices and the winding and core temperatures of power inductors. This allows the calculation results of dynamic load response degradation to eliminate interference factors introduced by operating temperature fluctuations, improving the accuracy of health characteristic indicators under variable temperature conditions. Furthermore, by weighted fusion of dynamic response parameters from different dimensions, the dynamic load response degradation can more comprehensively represent the degree of regulation capability degradation of the power module under transient conditions.

[0093] In one possible implementation of this application embodiment, when calculating the phase current imbalance of the power module, it can be implemented in the following ways, but not limited to: in response to the power module being a power module using a multi-phase parallel topology, within a preset time window, the current of each phase of the power module is sampled synchronously, and the average value of each phase current is calculated; the root mean square deviation of each phase current relative to the average value is calculated to obtain the phase current imbalance.

[0094] In the embodiments of this application, multiphase parallel topology refers to a circuit topology in which the output stage of a power module is composed of two or more power conversion phase branches with the same structure and connected in an interleaved parallel manner.

[0095] The preset time window refers to a pre-defined continuous time period used to determine the duration of current sampling data used in the calculation of phase-to-phase current imbalance. Since the currents of each phase have different phases in the interleaved parallel operation mode, the instantaneous values ​​of the currents of each phase are not the same at the same moment. Therefore, only by synchronously sampling the currents of each phase at the same moment can the obtained instantaneous current values ​​correspond to the same switching state. Only then can the calculated current differences truly reflect the differences in physical characteristics between the phase branches.

[0096] After obtaining the current of each phase, the arithmetic mean of multiple current samples for each phase within a preset time window is calculated to obtain the average current value of that phase within the preset time window. For example, if the power supply module is a three-phase parallel topology, the average current of the first phase, the average current of the second phase, and the average current of the third phase are calculated separately.

[0097] When calculating the phase-to-phase current imbalance, the average value of the current across all phases is first calculated. This average value is the sum of the average current values ​​of each phase divided by the number of phases. This average value represents the ideal balanced distribution of current across phases within a preset time window. Next, the deviation of the average current value of each phase from this average value is calculated; this is the difference between the current of each phase and the average value. This difference can be positive or negative; a positive value indicates that the current of that phase is higher than the average, and a negative value indicates that the current of that phase is lower than the average. The deviation values ​​of each phase are squared, summed, and then divided by the number of phases. Finally, the square root of the quotient is taken, and the result is the root mean square deviation (RMS deviation). This RMS deviation represents the phase-to-phase current imbalance. The larger the RMS deviation, the more uneven the current distribution across phases, and the more significant the characteristic differences between the power branches of each phase within the power module. The closer the RMS deviation is to zero, the more balanced the current distribution across phases, and the better the health of the power module.

[0098] This application only performs phase-to-phase current imbalance calculations for power modules employing multi-phase parallel topology, thus avoiding wasted computational resources. By synchronously sampling the current of each phase within a preset time window, temporal consistency of the phase current data is ensured. Using root mean square deviation to quantify the degree of imbalance between phase currents provides a comprehensive reflection of the current distribution across all phases, is more sensitive to minor imbalances, and can detect subtle differences between phases earlier, thereby providing information for health index assessment.

[0099] In one possible implementation of this application embodiment, the lightweight health index assessment model can also be obtained using, but is not limited to, the following methods: based on the computing power and memory of the controller in the target device, the neural network model is subjected to structured pruning and integer quantization to obtain the pruned and quantized neural network model; based on the historical aging data of the power module, the pruned and quantized neural network model is trained to obtain the lightweight health index assessment model; wherein, the lightweight health index assessment model is a neural network model that fuses depthwise separable convolutional and gated recurrent units.

[0100] In the embodiments of this application, the controller refers to the embedded microcontroller in the power module used to execute the health management method. The computing power can be comprehensively characterized by parameters such as the controller's clock frequency, whether it includes a floating-point unit, support for single-cycle multiply-accumulate instructions, and the number of fixed-point operations per second. Its memory refers to the capacity of the controller's internal static random access memory (SRAM) and non-volatile memory (NRAM). The former determines the intermediate data cache space required for model runtime, while the latter determines the storage space required for model parameters and structural information. Since the controllers selected for different models and specifications of power modules differ in computing power and memory configuration, the degree of model compression needs to be determined based on the specific hardware resources of the target controller before model deployment.

[0101] Structured pruning refers to a model compression technique that removes redundant structural units from a neural network model at the granular level, according to a preset importance evaluation criterion, thereby reducing the number of model parameters and computational cost. Integer quantization refers to a model compression technique that converts floating-point weights and floating-point activation values ​​in a neural network model into a low-bit-width integer representation (e.g., mapping a 32-bit floating-point number to an 8-bit integer INT8).

[0102] Historical aging data refers to the historical operating data and corresponding health status labels recorded by the power module throughout its entire life cycle. It includes at least the multi-dimensional real-time operating parameters collected by the power module at each historical sampling time, the corresponding health characteristic indicators, and the real health index or health status label determined by the actual failure time.

[0103] The application utilizes controller-based computing power and memory for structured pruning and integer quantization, enabling a lightweight health index assessment model to adaptively compress based on different hardware platforms. This maximizes model efficiency while maintaining prediction accuracy, solving the problem of deploying conventional deep learning models on resource-constrained embedded controllers. By configuring the lightweight health index assessment model as a neural network structure fusing deep separable convolutions and gated recurrent units, the model can leverage the spatial feature extraction capabilities of deep separable convolutions to fuse relevant information between different health feature indicators, while also fully utilizing the temporal modeling capabilities of gated recurrent units to capture the evolution of health feature indicators over time. This improves the accuracy of health index estimation while reducing computational and storage costs.

[0104] In one possible implementation of this application embodiment, when predicting the decay trajectory of the current health index, it can be achieved in the following ways, but not limited to: using a preset Kalman filter algorithm to predict the future decay trajectory of the current health index to obtain the predicted health trajectory; wherein, the preset Kalman filter algorithm dynamically adjusts the covariance matrix of process noise and measurement noise according to the fluctuation characteristics of the sequence data.

[0105] In the embodiments of this application, Kalman filtering is an optimal estimation algorithm based on the state-space model. Its basic principle is to perform a recursive optimal estimation of the true state of the system by fusing the system's state transition model (i.e., the evolution law of the system state over time) and the observation model (i.e., the observed values ​​corresponding to the state variables) and taking into account the statistical characteristics of process noise and observation noise.

[0106] The process noise covariance matrix (usually denoted as Q) and the measurement noise covariance matrix (usually denoted as R) are parameters that determine the performance of the filter estimation. Process noise reflects the degree of uncertainty in the state transition model, that is, the magnitude of the deviation between the evolution of the system state between adjacent time steps and the established state equation; measurement noise reflects the degree of deviation between the observed value and the true state of the system, that is, the magnitude of the error between the health index value output by the lightweight health index assessment model and the true health index.

[0107] The volatility characteristics of sequence data refer to the degree of dispersion and drastic change of the sequence data on the time axis, which can be specifically described by statistical characteristics such as the variance of the sequence, the absolute value of adjacent differences, the autocorrelation coefficient, or the spectral energy distribution.

[0108] This application predicts the future decay trajectory of the health index using a pre-defined Kalman filter algorithm. Based on the current observed health index value and combined with the trend reflected in historical health index sequences, it predicts the future decay trajectory of the current health index. The resulting predicted health trajectory simultaneously includes historical degradation trend information and correction information based on the current health index. Furthermore, by dynamically adjusting the process noise covariance matrix and measurement noise covariance matrix according to the fluctuation characteristics of the sequence data, this application solves the problem of decreased prediction performance caused by model mismatch in fixed-parameter Kalman filtering under different degradation stages and load conditions of the power module, thus improving the accuracy of remaining service life prediction.

[0109] In one possible implementation of this application embodiment, after determining the remaining service life of the power module, the following method may also be used, but is not limited to: in response to the remaining service life being lower than a preset remaining service life threshold, generating an alarm signal and performing derating operation on the target device.

[0110] In the embodiments of this application, the preset remaining lifetime threshold is a pre-set critical value for the lifespan, used as the basis for determining whether the power module has entered the end of its lifespan and requires triggering early warning and protection actions. The specific value of this threshold can be comprehensively set according to the reliability requirements of the power module, the importance of the target device's services, and the operation and maintenance response time. For example, for critical business target devices, the preset remaining lifetime threshold can be set to 720 hours (i.e., 30 days); for non-critical business target devices, it can be set to 168 hours (i.e., 7 days) or a shorter time.

[0111] Alarm signals are warning messages used to communicate to the target device management system or maintenance personnel that a power module has reached the end of its lifespan and requires attention or intervention. For example, alarm signals may include, but are not limited to: light signals emitted by the power module's fault indicator light at a specific flashing frequency or color; alarm or email notifications sent to the maintenance management platform via the target device management network; alarm event logs displayed in the data center management interface; and audible alarm signals emitted by a buzzer or speaker.

[0112] Derating refers to a protective control strategy that actively limits or reduces the output power capability of a power module to reduce the current and thermal stress on the power devices inside the power module, thereby slowing down the degradation rate of the power module and extending its actual usability.

[0113] This application achieves a process from state awareness and lifespan prediction to decision execution and maintenance response by generating an alarm signal and performing derating operation on the equipment after determining the remaining useful life and responding when the remaining useful life falls below a preset remaining useful life threshold. The generation of the alarm signal allows maintenance personnel to receive clear maintenance prompts before the power module actually fails; the execution of derating operation proactively reduces the electrothermal stress on the power module during the maintenance window, slowing its approach to the failure threshold, providing maintenance personnel with more response time, improving the fault tolerance and maintenance safety of the equipment's power supply system at the end of the power module's lifespan, and reducing the probability of sudden downtime.

[0114] In one possible implementation of this application embodiment, when generating sequence data, the following method may be used, but is not limited to: storing the current health index in a preset circular buffer; the circular buffer is used to store historical health indices within a preset time period in the past to generate sequence data.

[0115] In the embodiments of this application, the circular buffer is a fixed-capacity storage area implemented using a circular queue data structure. The circular buffer has a preset storage capacity, which is measured by the number of health index data records that can be stored. For example, it can be set to store the most recent 100, 200, or 500 health index records.

[0116] The preset time period refers to a fixed time length preceding the current moment, determined by the capacity of the circular buffer and the sampling frequency of the health index. For example, if the health index is sampled hourly and the circular buffer has a capacity of 720 records, the data stored in the circular buffer covers the health index data for the past 720 hours (i.e., 30 days). As new health indices are continuously written, the data stored in the circular buffer always remains within the preset time period, and older data exceeding that period is automatically discarded.

[0117] Sequence data is a sequence of health index values ​​read from a circular buffer in chronological order and their corresponding times. This sequence reflects the dynamic evolution of the health index over a continuous time axis.

[0118] This application uses a circular buffer to store health index data. Leveraging the fixed capacity and circular overlay characteristics of the circular buffer, it eliminates the need for dynamic allocation or release of storage space, enabling continuous maintenance of a fixed-length health index time series with limited memory resources. Furthermore, because the circular buffer automatically discards older data exceeding a preset time period, the series data focuses on the recent trends in health index changes, effectively utilizing the latest health status information for lifespan prediction and avoiding the problem of delayed prediction results due to the use of outdated data.

[0119] In one possible implementation of this application embodiment, the following methods may also be used, but are not limited to: when the power module performs a power-on self-test or during an idle period, perform a standard efficiency test and a dynamic load response test to update the baseline efficiency and standard response curves; in response to the deviation between the remaining service life and the actual operation and maintenance record being greater than a preset deviation threshold, obtain the real-time operating data of the target device, and update the lightweight health index assessment model based on the real-time operating data.

[0120] In the embodiments of this application, power-on self-test (POST) refers to the initialization self-diagnostic program executed by the power module each time it switches from a power-off state to a power-on state, before it formally outputs power to the target device load. During the power-on self-test phase, the main circuit of the power module has not yet supplied power to the load and is in an unloaded or very lightly loaded state. Performing standard efficiency tests and dynamic load response tests at this time will not affect the continuity of the target device's services. Idle periods refer to the time periods during which the target device load served by the power module is in a low-load or near-unloaded state during normal operation.

[0121] Actual operation and maintenance records refer to the real maintenance and replacement history data of the power module from its commissioning to the present moment. It includes at least the time of each actual maintenance, the maintenance type (such as routine inspection, component replacement, whole machine replacement, etc.), and the corresponding actual failure time or replacement time of the power module.

[0122] The preset deviation threshold is a critical value set in advance to determine whether the deviation between the predicted remaining service life and the actual operation and maintenance records exceeds an acceptable range. When the difference between the predicted remaining service life and the actual service time of the power module recorded in the operation and maintenance records exceeds this preset deviation threshold, it indicates that the current lightweight health index assessment model is no longer accurate in fitting the degradation pattern of this type of power module or under this usage condition, the reliability of the prediction results decreases, and the model needs to be updated.

[0123] Real-time operating data includes, but is not limited to: all multi-dimensional real-time operating parameters of the power module during historical operation, corresponding health characteristic indicators, actual maintenance and replacement time, and other data records.

[0124] This application tracks changes in baseline parameters caused by non-aging factors (such as temperature sensor bias drift, slow changes in current sampling resistor value, long-term changes in ambient temperature, etc.) by periodically performing standard efficiency tests and dynamic load response tests during the power module's power-on self-test or idle periods. These changes are then incorporated into the update of the baseline values, thereby ensuring that the calculation baselines for efficiency degradation rate and dynamic load response deterioration degree match the current state of the actual power module. This avoids the calculated results of health characteristic indicators deviating from the true degree of degradation due to outdated baseline values.

[0125] This application utilizes a model update mechanism to enable the lightweight health index assessment model to adaptively adjust based on actual operation and maintenance feedback. This allows the consistency and accuracy between the predicted results and the actual operating conditions to continuously improve over time, thereby maintaining a high level of predictive performance throughout the entire lifecycle of the equipment power module.

[0126] In one possible implementation of this application embodiment, when determining the target failure time, it can be achieved in the following ways, but not limited to: determining the trajectory time in the predicted trajectory where the health index is equal to a preset failure index threshold; and determining the trajectory time where the health index is equal to the preset failure index threshold as the target failure time.

[0127] In the embodiments of this application, by determining the trajectory moment in the predicted trajectory where the health index equals a preset failure index threshold and defining that moment as the target failure moment, a precise mapping from the predicted trajectory curve to the single-point failure moment is achieved.

[0128] Embodiments of this application also provide a health management system for a power module. Figure 2 A schematic diagram of a health management system for a power module provided in this application is shown below. Figure 2 As shown, it includes: The parameter acquisition module is used to synchronously acquire real-time operating parameters reflecting the aging status of multiple designated monitoring points corresponding to the power module of the target device. The real-time operating parameters include at least the thermal data, electrical data and dynamic response data of the power module. The feature calculation module is used to calculate health feature indicators that characterize the aging characteristics of the power supply module based on real-time operating parameters. An embedded processing module is used to input health characteristic indicators into a pre-trained lightweight health index evaluation model to obtain the current health index of the power module at the current moment. The lifespan prediction module is used to predict the decline trajectory of the current health index based on the sequence data corresponding to the current health index, and obtain the predicted health trajectory. The sequence data is a health index sequence that is constructed based on the current health index and the historical health index, and the health index changes over time. The life prediction module is also used to compare the predicted trajectory with the preset failure index threshold to determine the target failure time, and to determine the time difference between the target failure time and the current time as the remaining life of the power module. The data communication interface is used to report the current health index and remaining service life to the controller or remote monitoring platform of the target device.

[0129] In one possible implementation of this application embodiment, the parameter acquisition module includes: A high-precision synchronous sampling circuit is used to synchronously sample the voltage and current across the power switching devices of the power module. At least one temperature sensor network is arranged on the windings, core, and heat dissipation locations of the power inductor of the power module. The micro-disturbance signal generation and detection circuit is used to inject a micro-amplitude AC disturbance signal of a preset frequency into the output filter capacitor of the power supply module and detect the response caused by the micro-amplitude AC disturbance signal.

[0130] For a description of the features in the embodiment corresponding to the health management system of the power module, please refer to the relevant description of the embodiment corresponding to the health management method of the power module, which will not be repeated here.

[0131] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the health management method for a power module.

[0132] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the power module health management method when it is run.

[0133] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0134] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the power module health management method.

[0135] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the power module health management method.

[0136] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.

[0137] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0138] The above provides a detailed description of a health management method and system for a power module provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A health management method for a power module, characterized in that, include: At multiple designated monitoring points corresponding to the power module of the target device, real-time operating parameters reflecting the aging status of the designated monitoring points are collected synchronously. The real-time operating parameters include at least the thermal data, electrical data and dynamic response data of the power module. Based on the real-time operating parameters, a health characteristic index characterizing the aging characteristics of the power module is calculated, and the health characteristic index is input into a pre-trained lightweight health index evaluation model to obtain the current health index of the power module at the current moment. The decay trajectory of the current health index is predicted based on the sequence data corresponding to the current health index to obtain the predicted health trajectory. The sequence data is a health index sequence that varies over time, constructed based on the current health index and historical health indices. The predicted trajectory is compared with a preset failure index threshold to determine the target failure time, and the time difference between the target failure time and the current time is determined as the remaining service life of the power module.

2. The health management method for a power module according to claim 1, characterized in that, The synchronous collection of real-time operating parameters reflecting the aging status of multiple designated monitoring points corresponding to the power module of the target device includes: When the power switching device is in a fully on state, the voltage and current across the power switching device are sampled simultaneously to obtain the on-state voltage drop of the power switching device. The multiple designated monitoring points include at least the power switching device, the output filter capacitor, and the power inductor. The temperature is obtained by a temperature sensor integrated into the power switch device package, or calculated by the on-state voltage drop and a thermal resistance model that includes known temperature characteristics. A preset frequency micro-amplitude AC disturbance signal is injected into the output terminal of the power module, and the equivalent series resistance of the output filter capacitor is calculated based on the phase difference and amplitude ratio of the output voltage ripple and the output current ripple caused by the micro-amplitude AC disturbance signal. The electrical data includes at least the on-state voltage drop and junction temperature of the power switching device, the equivalent series resistance of the output filter capacitor, the output voltage ripple, and the output current ripple.

3. The health management method for a power module according to claim 2, characterized in that, The thermal data includes at least the winding temperature and core temperature of the power inductor, and the dynamic response data includes at least the ambient temperature and load current change rate of the power module.

4. The health management method for a power module according to claim 2, characterized in that, The health characteristic indicators that characterize the aging features of the power module based on the real-time operating parameters include: Based on the input voltage, input current, output voltage, and output current in the real-time operating parameters, the efficiency degradation rate of the power supply module is calculated. Based on the equivalent series resistance and current ripple of the output filter capacitor in the real-time operating parameters, the output voltage ripple increment of the power supply module is calculated. Based on the on-state voltage drop and junction temperature of the power switching device in the real-time operating parameters, as well as the winding temperature and core temperature of the power inductor, the dynamic load response degradation degree of the power module is calculated. Based on the current data in the real-time operating parameters, the phase-to-phase current imbalance of the power supply module is calculated.

5. The health management method for a power module according to claim 4, characterized in that, The calculation of the efficiency degradation rate of the power supply module based on the input voltage, input current, output voltage, and output current in the real-time operating parameters includes: The current efficiency of the power supply module is calculated based on the input voltage, input current, output voltage, and output current in the real-time operating parameters. The current efficiency is compared with the reference efficiency of the power module to obtain the percentage decrease of the current efficiency relative to the reference efficiency, and the percentage decrease is determined as the efficiency degradation rate, wherein the reference efficiency is measured by the power module within a historical calibration period or obtained by a preset method.

6. The health management method for a power module according to claim 4, characterized in that, The calculation of the dynamic load response degradation of the power module based on the on-state voltage drop and junction temperature of the power switching device, and the winding temperature and core temperature of the power inductor in the real-time operating parameters includes: A preset step load change test is performed on the power module to obtain the overshoot amplitude, recovery time and steady-state error of the output voltage; Based on the on-state voltage drop and junction temperature of the power switching device, as well as the winding temperature and core temperature of the power inductor, temperature compensation corrections are applied to the overshoot amplitude, the recovery time, and the steady-state error to obtain the corrected response parameters. The overshoot amplitude, recovery time and steady-state error in the corrected response parameters are weighted and calculated based on preset weights to obtain the target response curve. The target response curve is compared with the preset standard response curve to obtain the dynamic load response degradation degree.

7. The health management method for a power module according to claim 4, characterized in that, The calculation of the phase-to-phase current imbalance of the power module based on the current data in the real-time operating parameters includes: In response to the fact that the power module is a power module with a multi-phase parallel topology, the current of each phase of the power module is sampled synchronously within a preset time window, and the average value of each phase current is calculated. The root mean square deviation of each phase current relative to the average value is calculated to obtain the phase current imbalance.

8. The health management method for a power module according to claim 1, characterized in that, The method further includes: Based on the computing power and memory of the controller in the target device, the neural network model is subjected to structured pruning and integer quantization to obtain the pruned and quantized neural network model. Based on the historical aging data of the power module, the pruned and quantized neural network model is trained to obtain the lightweight health index assessment model. The lightweight health index assessment model is a neural network model that integrates deep separable convolutions and gated recurrent units.

9. The health management method for a power module according to claim 1, characterized in that, The step of predicting the decay trajectory of the current health index based on the sequence data corresponding to the current health index, and obtaining the predicted health trajectory, includes: The predicted health trajectory is obtained by using a preset Kalman filter algorithm to predict the future decline trajectory of the current health index. The preset Kalman filter algorithm dynamically adjusts the covariance matrix of process noise and measurement noise based on the fluctuation characteristics of the sequence data.

10. The health management method for a power module according to claim 1, characterized in that, After determining the time difference between the target failure time and the current time as the remaining service life of the power module, the method further includes: In response to the remaining service life being lower than a preset remaining service life threshold, an alarm signal is generated and the target device is subjected to derated operation.

11. The health management method for a power module according to claim 1, characterized in that, The method further includes: The current health index is stored in a preset circular buffer; The circular buffer is used to store historical health indices within a preset time period to generate the sequence data.

12. The health management method for a power module according to claim 1, characterized in that, The method further includes: During the power-on self-test or idle period of the power module, standard efficiency test and dynamic load response test are performed to update the baseline efficiency and standard response curves. In response to the deviation between the remaining service life and the actual operation and maintenance records being greater than a preset deviation threshold, the real-time operating data of the target device is obtained, and the lightweight health index assessment model is updated based on the real-time operating data.

13. The health management method for a power module according to claim 1, characterized in that, The step of comparing the predicted trajectory with a preset failure index threshold to determine the target failure time includes: Determine the trajectory moment in the predicted trajectory when the health index equals the preset failure index threshold; The trajectory time when the health index equals the preset failure index threshold is determined as the target failure time.

14. A health management system for a power module, characterized in that, include: The parameter acquisition module is used to synchronously acquire real-time operating parameters reflecting the aging status of multiple designated monitoring points corresponding to the power module of the target device. The real-time operating parameters include at least the thermal data, electrical data and dynamic response data of the power module. The feature calculation module is used to calculate health feature indicators that characterize the aging characteristics of the power supply module based on the real-time operating parameters. An embedded processing module is used to input the health characteristic indicators into a pre-trained lightweight health index evaluation model to obtain the current health index of the power module at the current moment. The lifespan prediction module is used to predict the decline trajectory of the current health index based on the sequence data corresponding to the current health index, and obtain the predicted health trajectory. The sequence data is a health index sequence that is constructed based on the current health index and the historical health index, and the health index changes over time. The life prediction module is also used to compare the predicted trajectory with a preset failure index threshold to determine the target failure time, and to determine the time difference between the target failure time and the current time as the remaining life of the power module. A data communication interface is used to report the current health index and the remaining service life to the controller or remote monitoring platform of the target device.

15. The health management system for the power module according to claim 14, characterized in that, The parameter acquisition module includes: A high-precision synchronous sampling circuit is used to synchronously sample the voltage and current across the power switching devices of the power module; At least one temperature sensor network is arranged on the windings of the power inductor of the power module, the magnetic core, and the heat dissipation location of the power module. The micro-disturbance signal generation and detection circuit is used to inject a micro-amplitude AC disturbance signal of a preset frequency into the output filter capacitor of the power supply module and detect the response caused by the micro-amplitude AC disturbance signal.