An ANPC converter dynamic hybrid modulation method based on online cost function learning
By employing a dynamic hybrid modulation method based on online cost function learning, the ANPC converter achieves a dynamic balance between efficiency, thermal reliability, and power quality under complex operating conditions. This solves the problem that traditional modulation strategies cannot adapt to device aging and environmental changes, thereby improving the system's reliability and lifespan.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG ELECTRICAL & ELECTRICAL GROUP SCIENCE & TECHNOLOGY RESEARCH CO LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-26
AI Technical Summary
Existing modulation and control strategies for ANPC converters struggle to balance efficiency, thermal reliability, and power quality under complex, dynamic, or non-standard operating conditions. Traditional methods cannot adapt to device aging and environmental changes, resulting in limited system reliability and lifespan.
A dynamic hybrid modulation method based on online cost function learning is adopted. By collecting electrical and thermal parameters in real time, a cost function mapping library with multidimensional lookup table is constructed. Combining a two-layer decision mechanism and an online learning mechanism, the optimal modulation strategy is dynamically selected to ensure system security and efficiency.
It enables autonomous optimization and safe operation of the ANPC converter under all operating conditions, improves the overall performance and long-term operational stability of the system, and avoids device damage caused by thermal runaway or midpoint imbalance.
Smart Images

Figure CN122292832A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power electronics technology, specifically to a dynamic hybrid modulation method for a multilevel active neutral point clamp (ANPC) converter, which is suitable for medium and high power applications requiring high reliability, high power density, and full-condition optimization. Background Technology
[0002] ANPC converters, with their low harmonics, high efficiency, and four-quadrant operation capabilities, are widely used in medium-to-high power applications such as new energy power generation, energy storage, and motor drives. However, their complex topology places higher demands on modulation strategies, requiring dynamic trade-offs between efficiency, thermal equilibrium, midpoint voltage stability, and output power quality. Traditional methods often employ fixed rules or offline lookup table strategies, which, while simple to implement, cannot adapt to actual operating conditions such as device aging, changes in ambient temperature, or sudden load changes. This can easily lead to junction temperature exceeding limits, midpoint imbalance, or even system shutdown, making it difficult to meet the requirements for high-reliability operation.
[0003] Existing modulation and control strategies for ANPC converters struggle to balance efficiency, thermal reliability, and power quality under complex, dynamic, or non-standard operating conditions. Traditional modulation methods often employ fixed rules or offline lookup table strategies, failing to dynamically adjust based on real-time junction temperature, midpoint voltage fluctuations, and environmental changes. They also lack proactive adaptability to gradual risks such as device aging and heat dissipation deterioration. Existing advanced control methods, such as model predictive control, have high computational complexity, making it difficult to complete global strategy search under safety constraints within microsecond-level control cycles. Furthermore, they rely on accurate models, resulting in poor robustness in the event of parameter mismatch. Current intelligent modulation schemes based on static mapping libraries lack online learning and confidence assessment mechanisms, making it impossible to continuously correct prediction biases from actual operating data. Under long-term operation or extreme conditions, they are prone to strategy misjudgments, leading to safety hazards such as overheating or midpoint imbalance.
[0004] The aforementioned defects not only limit the overall performance of the ANPC converter across the entire operating range, but may also lead to damage to power devices due to thermal runaway or DC-side instability, seriously affecting the reliability and lifespan of the system, resulting in limited system performance or even increased operational risks. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a dynamic hybrid modulation method for ANPC converters based on online cost function learning. This method combines microsecond-level real-time control, multi-objective intelligent decision-making, and reliable online learning capabilities to achieve autonomous optimization and safe operation of ANPC converters under all operating conditions.
[0006] To solve the aforementioned technical problem, the technical solution adopted by this invention is: a dynamic hybrid modulation method for ANPC converters based on online cost function learning, comprising the following steps: S01. Real-time acquisition of electrical and thermal parameters, including the DC bus voltage V of the converter. dc Midpoint voltage V np and the three-phase output current i a i b i c Thermal parameters include radiator temperature T sink and ambient temperature T; S02. Calculate the junction temperature of each power device based on the collected data, and calculate the current operating condition feature vector V. current =[P,φ, |ΔV np |, max(ΔT j ]], P represents the instantaneous active power of the converter, φ represents the power factor angle of the converter, |ΔV np | represents the degree of voltage imbalance at the midpoint, T j ΔT represents the junction temperature of the j-th power device. j This indicates the junction temperature difference of each power device; S03. Based on the current operating condition feature vector, match the nearest operating condition point in the cost function mapping library and execute a two-layer decision-making mechanism. The cost function mapping library is organized in the form of a multi-dimensional lookup table, with grid points X=[P,φ,T] consisting of the output active power P, power factor angle φ, and ambient temperature T. Each grid point stores the hard constraint cost g corresponding to all candidate modulation strategies. hard Soft target cost g soft Expected optimal strategy S opt The system learns state parameters, filters safety strategies that meet hard constraints in the first layer, selects the safety strategy with the lowest soft target cost as the final modulation strategy in the second layer, obtains the carrier type, modulation wave function and switching sequence corresponding to the final modulation strategy, generates PWM drive signal, and controls the operation of each power device in the ANPC converter in real time, thereby realizing dynamic hybrid modulation. S04. Obtain actual operating data, calculate prediction error, dynamically update confidence level, and execute recursive least squares method with adaptive forgetting factor under reliable conditions to perform online correction of cost function mapping library.
[0007] Furthermore, an initial cost function mapping library is constructed as the foundation for subsequent online decision-making and learning. The construction process is as follows: The cost function mapping library is organized in the form of a multidimensional lookup table. Output active power P, power factor angle φ, and ambient temperature T are selected as physical quantities representing the converter operating conditions. A discrete grid is constructed within the operating range of the selected physical quantities. For each grid point X=[P,φ,T], all candidate modulation strategy sets S are traversed. i And for each pair (X,S) i Calculate the hierarchical cost function; The hierarchical cost function includes hard constraint cost function and soft constraint cost function. The hard constraint cost function is used to ensure system safety and is defined as follows: , Among them, the penalty for exceeding the junction temperature limit for: , Penalty for severe midpoint voltage imbalance for: in This represents the maximum junction temperature of all power devices. This indicates the highest safe junction temperature threshold that the power device is allowed to operate at. , , These represent the capacitor voltage value on the DC bus, the capacitor voltage value below the DC bus, and the maximum allowable midpoint voltage imbalance safety threshold, respectively. Only when g hard When = 0, strategy S i It is considered safe and feasible under this operating condition; The soft constraint cost function is: , w e w t w q g represents the initial weighting coefficients. efficiency g thermal g quality These are the normalized efficiency, thermal balance, and waveform quality cost terms, respectively. The normalized efficiency cost term is: , The normalized heat balance cost term is: , The waveform quality cost term is: , in , , , , , These represent switching losses, conduction losses, rated power, average junction temperature, maximum junction temperature difference, and total harmonic distortion, respectively, while N represents the number of power devices. At each operating point, g will be able to hard =0 and g soft The strategy with the minimum value is denoted as the expected optimal strategy S for this operating condition. opt and its corresponding g soft The value is taken as the expected cost function value F. expected Store it in the cost function mapping library.
[0008] Furthermore, the two-tier decision-making mechanism is as follows: The first layer uses hard constraint filtering, querying the g of all candidate strategies at the nearest working point. hard Value; if there exists a value satisfying g hard The strategy with a value of 0 constitutes the security policy set S. safe If S safe If empty, enter damage control mode and directly select g. hard The minimum strategy is taken as the final decision S final ; The second layer is soft objective optimization, in S safe Internally recalculate g for each strategy soft Select g soft The minimum strategy is taken as the final decision S final .
[0009] Furthermore, step S04 specifically includes: Obtain the actual operating metrics from the previous dynamic update confidence period, including average switching loss, conduction loss, measured junction temperature fluctuation, midpoint voltage deviation, and THD. Calculate the actual cost function value F based on the acquired data. actual If no hard constraints are triggered during operation, that is, the actual maximum junction temperature does not exceed the safety limit T. j,max And the midpoint voltage imbalance |V c1 -V c2 | Always less than the allowable threshold V balance,max Then F actual This is equivalent to the soft target cost function value g recalculated based on measured data. soft Otherwise, it is judged as an abnormal working condition and F is given. actual Assign a high penalty value F penalty Then query the currently executed policy-condition pair (X). match ,S final Find the corresponding confidence score c in the cost function mapping library and calculate the prediction error: , The confidence level is dynamically updated based on the magnitude of the error: if ΔF < θ low This indicates that the prediction is reliable, and the confidence level is increased to c = min(c max ,c+1); if ΔF>θ high If the data is abnormal, the confidence level is significantly reduced to c = max(c min ,c−2), otherwise remain unchanged; θ low θ high These are the lower and upper thresholds for confidence level updates, respectively; only if the updated confidence level c > c trustOnly then is the learning update of the cost function mapping library triggered, c trust To update the confidence threshold, a recursive least squares method with an adaptive forgetting factor is used to learn and update the cost function mapping library.
[0010] Furthermore, the learning and updating process of the cost function mapping library is as follows: Designing forgetting factors for: , Where λ base The basal forgetting factor is given, and τ is the sensitivity coefficient. Perform RLS iteration: , , , , Using the covariance matrix and expected cost function value from the previous iteration, the first iteration uses the initial covariance matrix and initial expected cost function value. , The covariance matrix and expected cost function value for this iteration; Updated F expected P(k) and P(k) are written back to the corresponding entries in the cost function mapping library.
[0011] Furthermore, θ low =0.05, θ high =0.2, c trust = 2.0.
[0012] Furthermore, a strategy identifier is set for each modulation strategy, and the corresponding carrier type, modulation waveform function and switching sequence are called from the PWM parameter library based on the strategy identifier to generate the PWM drive signal.
[0013] Furthermore, steps S01 and S02 are executed in microsecond cycles, and step S04 is executed in millisecond to second cycles.
[0014] The present invention also discloses an ANPC converter dynamic hybrid modulation system based on online cost function learning, including a multi-dimensional sensor array, a real-time controller, an AI coprocessor, and a general-purpose processor; The multi-dimensional sensor array includes electrical sensors for acquiring the converter's electrical parameters and thermal sensors for acquiring thermal parameters, including the converter's DC bus voltage V. dc Midpoint voltage V np and the three-phase output current i a i b ic Thermal parameters include radiator temperature T sink Both the electrical and thermal sensors are connected to the signal conditioning circuit, along with the ambient temperature T. The real-time controller is connected to the signal conditioning circuit to synchronously acquire various parameters output by the multi-dimensional sensor array, calculate the junction temperature of each power device, and calculate the current operating condition feature vector V. current =[P, φ, |ΔV np |, max(ΔT j ]], P represents the instantaneous active power of the converter, φ represents the power factor angle of the converter, |ΔV np | represents the degree of voltage imbalance at the midpoint, T j ΔT represents the junction temperature of the j-th power device. j This indicates the junction temperature difference of each power device; The AI coprocessor is connected to the real-time controller. Based on the current operating condition feature vector, it matches the nearest operating condition point in the cost function mapping library and executes a two-layer decision-making mechanism. The cost function mapping library is organized in the form of a multi-dimensional lookup table, with grid points X=[P,φ,T] consisting of the output active power P, power factor angle φ, and ambient temperature T. Each grid point stores the hard constraint cost g corresponding to all candidate modulation strategies. hard Soft target cost g soft Expected optimal strategy S opt The learning state parameters are used to first screen for safety policies that meet hard constraints, and then the second layer selects the safety policy with the lowest soft target cost as the final modulation policy. At the same time, the AI coprocessor calculates the prediction error based on the actual running data provided by the general processor, dynamically updates the confidence level, and performs online correction of the cost function mapping library by recursive least squares method with adaptive forgetting factor under credible conditions. The final modulation strategy is returned to the real-time controller, which calls the carrier type, modulation function and switching sequence corresponding to the final modulation strategy to generate a PWM drive signal and send it to the ANPC converter. The general-purpose processor is connected to the real-time controller, AI coprocessor, and external non-volatile memory. It is responsible for system initialization, non-volatile memory access, mapping library persistence, and high-level task scheduling. It shares the mapping library storage area with the AI coprocessor and periodically extracts actual running metrics from the historical data buffer of the real-time controller to trigger the online learning process. The general-purpose processor loads the initial cost function mapping library through non-volatile memory access and saves the updated learning state during operation.
[0015] Furthermore, the real-time controller, AI coprocessor, and general-purpose processor are deployed on the ZYNQ heterogeneous computing platform. The real-time controller is deployed in programmable logic or a real-time microcontroller, and synchronously acquires various acquisition parameters output by the multi-dimensional sensor array through a multi-channel ADC interface. The AI coprocessor is integrated into the programmable logic, and the general-purpose processor is an ARM Cortex-A series application processor.
[0016] The beneficial effects of this invention are as follows: The proposed dynamic hybrid modulation method and system for ANPC converters based on online cost function learning, compared with existing ANPC modulation technologies, constructs hierarchical cost functions for fusion efficiency, thermal balance, and power quality. Combined with real-time operating condition perception and online learning mechanisms, the system can dynamically select the optimal modulation strategy under complex scenarios such as load fluctuations, high-temperature aging, and heat dissipation deterioration, significantly enhancing its all-condition adaptive optimization capability. A two-layer decision-making mechanism of "hard constraint priority, soft objective suboptimal" ensures that device safety limits are not violated under any operating condition. A confidence-gated online learning mechanism is introduced to guarantee long-term robustness and stability. Through a recursive least squares (RLS) algorithm with an adaptive forgetting factor, and by periodically persisting the learning results (including the covariance matrix and confidence level) to non-volatile memory, continuous evolution and power-off memory are supported. Attached Figure Description
[0017] Figure 1 This is a block diagram of the modulation system described in Example 1; Figure 2 Detailed diagram of heterogeneous computing platform architecture; Figure 3 A flowchart of the online intelligent modulation system; Figure 4 This is a schematic diagram of the hierarchical cost function structure; Figure 5 Waveform diagram of modulation strategy switching; Figure 6 Junction temperature diagram of the modulation strategy switching device. Detailed Implementation
[0018] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0019] Example 1 This embodiment discloses a dynamic hybrid modulation system for ANPC converters based on online cost function learning. The system comprises a multi-dimensional sensor array, a general-purpose processor (GP), a real-time control unit (RCU), and an AI processing unit (APU). These four components work collaboratively via on-chip shared memory and a high-speed interconnect bus, and are deployed on a heterogeneous computing hardware platform, Xilinx Zynq SoC. The interconnections between the modules are as follows: Figure 1 , 2 As shown.
[0020] The multi-dimensional sensor array is used to comprehensively perceive the operating status and environmental conditions of the ANPC converter, including electrical sensors composed of high-precision isolated voltage and current sensors, for real-time acquisition of the DC bus voltage V. dc Midpoint voltage V np and the three-phase output current i a i b i c Key electrical parameters; distributed temperature sensors, such as thermal sensors composed of digital temperature chips or thermocouples, are arranged near the power device substrate, heat sink, and equipment housing to monitor the heat sink temperature T. sink The ambient temperature T provides data support for junction temperature estimation and thermal safety assessment; all sensor signals are electrically isolated before being connected to the analog or digital input interface of the real-time controller (RCU) to ensure reliable sampling and safe operation of the system in high-noise and high-common-mode voltage environments.
[0021] The general-purpose processor (GP) is an ARM Cortex-A series application processor running an embedded operating system. It is responsible for system initialization, non-volatile memory access, mapping library persistence, and high-level task scheduling. It shares the mapping library storage area with the APU via an AXI high-speed interface and periodically retrieves actual runtime metrics from the RCU's historical data buffer to trigger the online learning process. The GP also connects to external non-volatile memory via QSPI or eMMC interfaces to load the initial cost function mapping library and save the updated learning state during runtime.
[0022] The Real-Time Controller (RCU) is deployed in a programmable logic unit (PL) or a real-time microcontroller, executing low-level control tasks in microsecond cycles. It synchronously acquires various parameters from the multi-dimensional sensor array via a high-precision multi-channel ADC interface. Based on a pre-calibrated switching loss and thermal network model, it estimates the junction temperature of each power device in real time and calculates the current operating condition feature vector V. current =[P, φ, |ΔV np |, max(ΔTj )], P represents instantaneous active power, φ represents the power factor angle, |ΔV np | represents the voltage imbalance at the midpoint, max(ΔT) j The vector represents the maximum junction temperature difference and is sent to the APU via the on-chip AXI-Stream bus. At the same time, the RCU, based on the strategy identifier issued by the APU, calls the corresponding carrier type, modulation waveform and switching sequence from the internal PWM parameter library to generate a high-resolution PWM signal to drive the ANPC converter power devices.
[0023] The AI coprocessor (APU) is integrated into the programmable logic (PL) and includes a condition matching module, a hierarchical decision-making module, a parallel cost function evaluation module, and an improved RLS learning acceleration module. After receiving the feature vector sent by the RCU, the APU matches the nearest operating point in the cached cost function mapping library and performs two-layer decision-making: the first layer filters the set of strategies that meet hard constraints (junction temperature and midpoint voltage safety), and the second layer selects the strategy with the lowest soft target cost (efficiency, thermal equilibrium, power quality) as the final modulation strategy. Simultaneously, based on the actual operating data provided by the GP, the APU calculates the prediction error, dynamically updates the confidence level, and performs online correction of the mapping library using recursive least squares (RLS) with an adaptive forgetting factor under reliable conditions. The updated mapping library entries (including F...) expected The covariance matrix P(k) and confidence level c) reside continuously in the PL shared memory and are periodically written to non-volatile memory by GP to achieve power-off memory.
[0024] The cost function mapping library is organized in the form of a multidimensional lookup table. The operating condition dimension includes the output active power P, the power factor angle φ, and the ambient temperature T. Each grid point stores the hard constraint cost g corresponding to all candidate modulation strategies (such as LS-PWM, PS-PWM, and hybrid strategies). hard Soft target cost g soft Expected optimal strategy S opt and learning state parameters.
[0025] Through the above architecture, this invention realizes closed-loop intelligent modulation of perception-decision-execution-learning, enabling the ANPC converter to autonomously balance efficiency, thermal reliability and output power quality across the entire operating range, significantly improving the overall system performance and long-term operational safety.
[0026] Example 2 This embodiment discloses a dynamic hybrid modulation method for ANPC converters based on online cost function learning, realizing dynamic hybrid modulation of ANPC converters, such as... Figure 3 As shown, it includes the following steps: Step 1: Build the initial cost function mapping library offline.
[0027] Before the system's initial operation, an initial cost function mapping library needs to be constructed through offline simulation and experiments, serving as the foundation for subsequent online decision-making and learning. This mapping library is organized in the form of a multidimensional lookup table, with its operating condition dimension selecting the three most representative physical quantities in converter operation: output active power P, power factor angle φ, and ambient temperature T, and constructing a discrete grid within its typical operating range. For each grid point, all candidate modulation strategy sets Si are traversed, and a hierarchical cost function is calculated for each pair (X, Si). For example... Figure 4 As shown, the hard constraint cost function is used to ensure system safety and is defined as follows: , The penalty for exceeding the junction temperature limit is as follows: , The penalty for severe voltage imbalance at the midpoint is: , Where g hard For hard constraint cost function, This is a penalty item for exceeding the junction temperature limit. This is a penalty term for severe imbalance in the midpoint voltage. This represents the maximum junction temperature of all power devices. This indicates the highest safe junction temperature threshold that the power device is allowed to operate at. , , These represent the capacitor voltage values on the DC bus, the capacitor voltage values off the DC bus, and the maximum permissible midpoint voltage imbalance safety threshold, respectively. Only when g hard When = 0, strategy Si is considered safe and feasible under this condition. Based on this, the soft objective cost function is further evaluated for all safe strategies: , w e ,w t ,w q g represents the initial weighting coefficients. efficiency g thermal g quality These represent the normalized efficiency, thermal balance, and waveform quality cost terms, respectively. The efficiency cost term normalizes the total loss as follows: , The normalized thermal imbalance cost term reflects the junction temperature difference between the switching transistors: , The waveform quality cost can be represented by the total harmonic distortion (THD) of the output current: , in This is a penalty item for exceeding the junction temperature limit. This is a penalty term for severe imbalance in the midpoint voltage. This represents the maximum junction temperature of all power devices. This indicates the highest safe junction temperature threshold that the power device is allowed to operate at. , , These represent the capacitor voltage value on the DC bus, the capacitor voltage value below the DC bus, and the maximum allowable midpoint voltage imbalance safety threshold, respectively.
[0028] At each operating point, g will be able to hard =0 and g soft The strategy with the minimum value is marked as the expected optimal strategy for this working condition, and its corresponding g is assigned to it. soft The value is taken as the expected cost function value F. expected Store in the mapping library.
[0029] Figure 5 , Figure 6 To facilitate offline simulation before system operation, an initial cost function mapping library was constructed through simulation and bench testing, serving as the foundation for subsequent online decision-making and learning. Simultaneously, through... Figure 5 , Figure 6 Simulation results also show that switching modulation strategies affects current, junction temperature, etc., which illustrates the necessity of using this patent for dynamic hybrid modulation under complex operating conditions.
[0030] Step 2: Hardware platform and algorithm initialization.
[0031] After the system powers on, the GP loads the initial mapping library from non-volatile memory to the APU's high-speed shared memory; simultaneously, the RCU completes the initialization configuration of peripherals such as ADC and PWM; the APU then initializes the online learning parameters for each entry in the mapping library, including the initial covariance matrix P(0)=αI of recursive least squares (RLS) and the confidence score. C = C min Where α is the initialization coefficient, taking a large positive real number, representing the algorithm's initial guess value (i.e., F obtained from offline simulation). expected The degree of trust in ), where I is the identity matrix.
[0032] Step 3: Real-time state awareness and feature extraction.
[0033] The real-time dynamic modulation cycle operates on a microsecond-level timescale, with the RCU (Real-Time Control Unit) handling data acquisition and the APU (Action Processing Unit) responsible for intelligent decision-making. Within each control cycle, the RCU acquires real-time sensing data from a multi-dimensional sensor array, including the three-phase output current i. a i b ic DC bus voltage V dc Midpoint voltage V np and radiator temperature T sink Based on these raw signals, combined with a pre-calibrated switching loss model and a simplified thermal network model, the RCU estimates the junction temperature T of the six power switches in real time. j1 ~Tj6, and calculate the key characteristics of the current operating state: instantaneous active power P, power factor angle φ, and neutral point voltage imbalance |ΔV. np |=|V c1 -V c2 | and the maximum junction temperature difference max(ΔT) j These features are encapsulated as state vectors: .
[0034] Step 4: Real-time modulation strategy selection based on hierarchical decision-making.
[0035] The packaged state vector is sent to the APU via the on-chip high-speed bus. After receiving the vector, the APU matches the nearest reference operating point X in the mapping library using a weighted Euclidean distance. match Then a two-tier decision-making mechanism is implemented.
[0036] The first layer is a hard constraint filter: query X match g of all candidate strategies hard Value; if there exists a value satisfying g hard The strategy with a value of 0 constitutes the security policy set S. safe If S safe If empty, enter damage control mode and directly select g. hard The minimum strategy is taken as the final decision S final And skip the performance optimization steps.
[0037] The second layer is soft objective optimization: in S safe Internally recalculate g for each strategy soft Its weighting coefficient can be dynamically adjusted according to the application scenario, selecting the factor that makes g... soft The minimum strategy is S final The APU then sends the strategy identifier to the RCU, which uses its internal pre-stored PWM parameter library to call the corresponding carrier type, modulation waveform function, and switching sequence to generate a high-precision PWM drive signal. This signal controls the operation of each power device in the ANPC converter in real time, thereby achieving dynamic hybrid modulation.
[0038] Step 5: Online learning and mapping library update cycle.
[0039] On a timescale of milliseconds to seconds, such as every 100ms, the GP and APU collaboratively perform system performance evaluation. The GP first extracts actual operating metrics from the historical data buffer maintained by the RCU for the previous evaluation period, including key parameters such as average switching loss, conduction loss, measured junction temperature fluctuation, midpoint voltage deviation, and THD. After receiving this data, the APU calculates the actual cost function value F. actual : , g soft_actul Substitute the measured value at the current time into the soft objective function g_soft.
[0040] As can be seen from the above formula, if no hard constraints are triggered during operation, that is, the actual maximum junction temperature does not exceed the safety limit T. j,max And the midpoint voltage imbalance |Vc1−Vc2| is always less than the allowable threshold V. balance,max ), then F actua l is equal to the soft target cost function value g recalculated based on measured data. soft Otherwise, it is judged as an abnormal operating condition and a high penalty value F is assigned. penalty (For example, set it to 10 times the upper limit of the normal cost range). Then, the APU queries the currently executed policy-condition pair (X). match ,S final Find the corresponding confidence score c in the mapping library and calculate the prediction error: , The confidence level is dynamically updated based on the magnitude of the error: if ΔF < θ low (e.g., 0.05) indicates that the prediction is reliable, and the confidence level is increased to c = min(c max ,c+1); if ΔF>θ high If the value is 0.2, then the data is considered abnormal, and the confidence level is significantly reduced to c = max(c min (c-2); otherwise, remain unchanged. Only if the updated confidence level c > c trust (e.g., version 2.0) The learning update of the mapping library is only triggered then. At this time, the APU starts a dedicated computing unit to execute recursive least squares (RLS) with an adaptive forgetting factor. The forgetting factor is designed as follows: , Where λ base =0.95 is the basic forgetting factor, and τ is the sensitivity coefficient (e.g., 0.1). This design allows λ to approach λ when the prediction error is small. base The update is more sensitive; when the error is large, λ approaches 1, and the update tends to be conservative. Then, RLS iteration is performed: , , ; in , Given the covariance matrix and expected cost function value from the previous iteration, in the first iteration... Using the initial covariance matrix, F in the first iteration expected (k) uses the expected cost function value stored when the initial cost function mapping library is built offline. This function value is the theoretical optimal value obtained from offline data through a large number of offline simulations and bench tests before the system's first run. , These are the covariance matrix and expected cost function values for this iteration.
[0041] Updated F expected P(k) and P(k) are written back to the corresponding entries in the mapping library. Finally, the GP periodically, such as once per second, or before the system is safely shut down, persists the complete mapping library and its learning state in the APU shared memory, including the covariance matrix P(k) and confidence level c of each entry, to non-volatile memory, thereby realizing the long-term accumulation of operating experience and the function of power-off memory.
[0042] Through the closed-loop coordinated operation of the above five steps, this invention realizes the leap from "fixed rules" to "autonomous intelligence" in the modulation strategy of ANPC converter. It can dynamically and adaptively maintain the optimal balance between efficiency, thermal reliability and output power quality across the entire operating range, significantly improving the overall system performance and environmental adaptability.
[0043] The above embodiments are merely illustrative of the technical solutions and implementation processes of the present invention and are not intended to limit the scope of protection of the present invention. Any technical solutions implemented through equivalent substitutions, adaptive modifications, or extended applications based on the system architecture, control method, and learning mechanism described in the present invention are all within the scope of protection of the present invention.
Claims
1. A dynamic hybrid modulation method for ANPC converters based on online cost function learning, characterized in that: Includes the following steps: S01. Real-time acquisition of electrical and thermal parameters, including the DC bus voltage V of the converter. dc Midpoint voltage V np and the three-phase output current i a i b i c Thermal parameters include radiator temperature T sink and ambient temperature T; S02. Calculate the junction temperature of each power device based on the collected data, and calculate the current operating condition feature vector V. current =[P, φ, |ΔV np |, max(ΔT j ]], P represents the instantaneous active power of the converter, φ represents the power factor angle of the converter, |ΔV np | represents the degree of voltage imbalance at the midpoint, T j ΔT represents the junction temperature of the j-th power device. j This indicates the junction temperature difference of each power device; S03. Based on the current operating condition feature vector, match the nearest operating condition point in the cost function mapping library and execute a two-layer decision-making mechanism. The cost function mapping library is organized in the form of a multi-dimensional lookup table, with grid points X=[P,φ,T] consisting of the output active power P, power factor angle φ, and ambient temperature T. Each grid point stores the hard constraint cost g corresponding to all candidate modulation strategies. hard Soft target cost g soft Expected optimal strategy S opt The system learns state parameters, filters safety strategies that meet hard constraints in the first layer, selects the safety strategy with the lowest soft target cost as the final modulation strategy in the second layer, obtains the carrier type, modulation wave function and switching sequence corresponding to the final modulation strategy, generates PWM drive signal, and controls the operation of each power device in the ANPC converter in real time, thereby realizing dynamic hybrid modulation. S04. Obtain actual operating data, calculate prediction error, dynamically update confidence level, and execute recursive least squares method with adaptive forgetting factor under reliable conditions to perform online correction of cost function mapping library.
2. The dynamic hybrid modulation method for ANPC converter based on online cost function learning according to claim 1, characterized in that: An initial cost function mapping library is constructed as the foundation for subsequent online decision-making and learning. The construction process is as follows: The cost function mapping library is organized in the form of a multidimensional lookup table. Output active power P, power factor angle φ, and ambient temperature T are selected as physical quantities representing the converter operating conditions. A discrete grid is constructed within the operating range of the selected physical quantities. For each grid point X=[P,φ,T], all candidate modulation strategy sets S are traversed. i And for each pair (X,S) i Calculate the hierarchical cost function; The hierarchical cost function includes hard constraint cost function and soft constraint cost function. The hard constraint cost function is used to ensure system safety and is defined as follows: , Among them, the penalty for exceeding the junction temperature limit for: , Penalty for severe midpoint voltage imbalance for: , in This represents the maximum junction temperature of all power devices. This indicates the highest safe junction temperature threshold that the power device is allowed to operate at. , , These represent the capacitor voltage value on the DC bus, the capacitor voltage value below the DC bus, and the maximum allowable midpoint voltage imbalance safety threshold, respectively. Only when g hard When = 0, strategy S i It is considered safe and feasible under this operating condition; The soft constraint cost function is: , w e w t w q g represents the initial weighting coefficients. efficiency g thermal g quality These are the normalized efficiency, thermal balance, and waveform quality cost terms, respectively. The normalized efficiency cost term is: , The normalized heat balance cost term is: , The waveform quality cost term is: , in , , , , , These represent switching losses, conduction losses, rated power, average junction temperature, maximum junction temperature difference, and total harmonic distortion, respectively, while N represents the number of power devices. At each operating point, g will be able to hard =0 and g soft The strategy with the minimum value is denoted as the expected optimal strategy S for this operating condition. opt and its corresponding g soft The value is taken as the expected cost function value F. expected Store it in the cost function mapping library.
3. The dynamic hybrid modulation method for ANPC converter based on online cost function learning according to claim 2, characterized in that: The two-tier decision-making mechanism is as follows: The first layer uses hard constraint filtering, querying the g of all candidate strategies at the nearest working point. hard Value; if there exists a value satisfying g hard The strategy with a value of 0 constitutes the security policy set S. safe If S safe If empty, enter damage control mode and directly select g. hard The minimum strategy is taken as the final decision S final ; The second layer is soft objective optimization, in S safe Internally recalculate g for each strategy soft Select g soft The minimum strategy is taken as the final decision S final .
4. The dynamic hybrid modulation method for ANPC converter based on online cost function learning according to claim 1, characterized in that: Step S04 is as follows: Obtain the actual operating metrics from the previous dynamic update confidence period, including average switching loss, conduction loss, measured junction temperature fluctuation, midpoint voltage deviation, and THD. Calculate the actual cost function value F based on the acquired data. actual , , g soft_actul Substituting the measured value at the current moment into the soft objective function g_soft, if no hard constraints are triggered during the operation, that is, the actual maximum junction temperature does not exceed the safety limit T. j,max And the midpoint voltage imbalance |V c1 -V c2 | Always less than the allowable threshold V balance,max Then F actual This is equivalent to the soft target cost function value g recalculated based on measured data. soft Otherwise, it is judged as an abnormal working condition and F is given. actual Assign a high penalty value F penalty Then query the currently executed policy-condition pair (X) match ,S final Find the corresponding confidence score c in the cost function mapping library and calculate the prediction error: , The confidence level is dynamically updated based on the magnitude of the error: if ΔF < θ low This indicates that the prediction is reliable, and the confidence level is increased to c = min(c max ,c+1); if ΔF>θ high If the data is abnormal, the confidence level is significantly reduced to c = max(c min ,c−2), otherwise remain unchanged; θ low θ high These are the lower and upper thresholds for confidence level updates, respectively; only if the updated confidence level c > c trust Only then is the learning update of the cost function mapping library triggered, c trust To update the confidence threshold, a recursive least squares method with an adaptive forgetting factor is used to learn and update the cost function mapping library.
5. The dynamic hybrid modulation method for ANPC converter based on online cost function learning according to claim 4, characterized in that: The learning and updating process of the cost function mapping library is as follows: Designing forgetting factors for: , Where λ base The basal forgetting factor is given, and τ is the sensitivity coefficient. Perform RLS iteration: , , , , Using the covariance matrix and expected cost function value from the previous iteration, the first iteration uses the initial covariance matrix and initial expected cost function value. , The covariance matrix and expected cost function value for this iteration; Updated F expected P(k) and P(k) are written back to the corresponding entries in the cost function mapping library.
6. The dynamic hybrid modulation method for ANPC converter based on online cost function learning according to claim 4, characterized in that: i low =0.05,θ high =0.2,c trust = 2.0。 7. The dynamic hybrid modulation method for ANPC converter based on online cost function learning according to claim 1, characterized in that: For each modulation strategy, a strategy identifier is set. Based on the strategy identifier, the corresponding carrier type, modulation waveform function and switching sequence are called from the PWM parameter library to generate the PWM drive signal.
8. The dynamic hybrid modulation method for ANPC converter based on online cost function learning according to claim 1, characterized in that: Steps S01 and S02 run in microsecond cycles, while step S04 runs in millisecond to second cycles.