Reinforcement learning sequential MPC based multilevel inverter control method and system

By constructing a three-layer optimization structure using a reinforcement learning-based sequential MPC method, and dynamically adjusting the tolerance boundary value, the problem of parameter tuning blind zone and target conflict in traditional multi-level inverter control is solved, thereby improving the dynamic response and adaptability of the system.

CN120880221BActive Publication Date: 2026-05-15SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511068033.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2026-05-15
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

In traditional multilevel inverter control, multi-objective optimization problems rely on empirical parameter settings, leading to a cumbersome tuning process, and in variable operating conditions, there are problems such as dynamic recovery hysteresis or performance degradation.

Method used

A sequential MPC method based on reinforcement learning is adopted to construct a three-layer optimization structure, namely current tracking, flying capacitor voltage balancing and midpoint potential voltage balancing. The tolerance boundary values ​​are autonomously learned through offline reinforcement learning, and the tolerance boundaries of each layer are dynamically adjusted to achieve independent optimization.

Benefits of technology

It enables automatic design and optimization of parameters in multi-objective systems, avoids objective conflicts in the traditional weighted summation method, and improves the dynamic response capability and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120880221B_ABST
    Figure CN120880221B_ABST
Patent Text Reader

Abstract

The application discloses a multi-level inverter control method and system based on reinforcement learning sequential MPC, and relates to the technical field of power electronics control, and comprises the following steps: a three-layer sequential optimization structure of current tracking-flying capacitor voltage balance-neutral point potential voltage balance is constructed in turn; a multi-objective reward function is designed by taking system state as observation input, taking current tracking tolerance boundary and flying capacitor voltage balance tolerance boundary as action output, so that dynamic balance of current tracking, flying capacitor voltage and neutral point voltage is realized; multiple control targets of optimized current tracking, flying capacitor voltage balance and neutral point potential voltage balance are optimized, and each layer is independently optimized and controlled; meanwhile, based on autonomous learning of offline reinforcement learning, optimization of multi-objective tolerance boundary values is realized, and the key problems of low artificial parameter adjustment efficiency and existing parameter setting blind area in a multi-objective system are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power electronic control technology, and in particular to a multilevel inverter control method and system based on reinforcement learning sequential MPC. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Multilevel inverters, as core devices in renewable energy conversion systems, need to simultaneously achieve multiple objectives, such as accurate output voltage tracking and capacitor voltage balance. Traditional model predictive control (MPC) uses weighting factors to combine multiple control objectives into a single value function for solution. By modifying the weighting factors, the emphasis of the control objectives can be adjusted. This multi-objective optimization problem has long faced difficulties in engineering practice: traditional weighted control methods rely on empirical parameter settings, and the mismatch of their physical dimensions leads to a lack of theoretical design criteria. Therefore, the appropriate values ​​of the weighting factors can only be determined through extensive experiments, rather than specific theories, resulting in a cumbersome tuning process.

[0004] To address this issue, a sequential model predictive control (MPC) approach is employed, partially resolving the aforementioned problems through a hierarchical optimization structure. In each optimization stage of the sequential MPC hierarchical optimization structure, a certain error range is allowed for the current optimization objective, quantified by a tolerance boundary.

[0005] However, the fixed tolerance boundary setting mode presents a contradiction in variable operating conditions: when the parameter setting is too small, the system adjustment margin may be too small, which may lead to a delay in the dynamic recovery process; when the parameter setting is too large, it will deteriorate the performance indicators of the upstream stage (such as harmonic distortion of the output current).

[0006] In the current field of power electronics control, multi-objective optimization scenarios lack dynamic parameter adjustment mechanisms. This inherent limitation of the control architecture has become a core bottleneck for upgrading power quality equipment to higher reliability and adaptability, and there is an urgent need to build innovative solutions with dynamic calculation capabilities. Summary of the Invention

[0007] To address the aforementioned issues, this invention proposes a multilevel inverter control method and system based on reinforcement learning sequential MPC. This method optimizes multiple control objectives, including current tracking, flying capacitor voltage balance, and midpoint potential voltage balance, with each layer performing independent optimization control. Furthermore, based on offline reinforcement learning, it achieves optimization of multi-objective tolerance boundary values, thus solving the key challenges of low efficiency in manual parameter tuning and the existence of parameter tuning blind spots in multi-objective systems.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] In a first aspect, the present invention provides a multilevel inverter control method based on reinforcement learning sequential MPC, comprising:

[0010] A three-layer sequential optimization structure is constructed, which progresses sequentially from current tracking to flying capacitor voltage balancing to midpoint potential voltage balancing. Each layer solves for the optimal value based on its own value function and selects a candidate vector set according to the tolerance boundary. Each layer optimizes within the candidate vector set selected by the upper layer to generate the optimal control command.

[0011] The observation inputs are current, flying capacitor voltage, midpoint potential voltage and corresponding deviation values. The action outputs are current tracking tolerance boundary and flying capacitor voltage balance tolerance boundary. A multi-objective reward function combining steady-state reward function and transient reward function is designed.

[0012] The reinforcement learning agent is trained offline using a multi-objective reward function. The trained reinforcement learning agent is then used to determine the optimal tolerance boundary values ​​for the current tracking layer and the flying capacitor balance layer based on real-time observations, thereby achieving dynamic balance between current tracking, flying capacitor voltage, and midpoint voltage.

[0013] As an alternative implementation, in the current tracking layer, a value function is designed based on the deviation between the reference current and the actual current. All candidate vectors are evaluated in this way, the minimum value function is determined, and a second set of candidate vectors within the allowable deviation range is generated based on the current tracking tolerance boundary.

[0014] In the flying capacitor voltage balancing layer, a value function is designed based on the deviation between the reference flying capacitor voltage and the actual flying capacitor voltage to determine the minimum value function in the second-layer candidate vector set, and a third-layer candidate vector set within the allowable deviation range is generated based on the flying capacitor voltage balancing tolerance boundary.

[0015] In the midpoint potential voltage balance layer, the optimal vector that maintains the stability of the DC midpoint potential voltage is selected from the candidate vector set in the third layer, thereby obtaining the optimal control command.

[0016] As an alternative implementation, in the current tracking layer, the designed value function for:

[0017] ;

[0018] The constraints satisfied by the second-level candidate vector set are: ;

[0019] in, For current tracking tolerance boundaries; and For reference current, and yes k The current at time +1 It is the minimum current value function;

[0020] The value function of the design across the capacitor voltage balance layer. for:

[0021] ;

[0022] The constraints satisfied by the third-level candidate vector set are: ;

[0023] in, To balance the tolerance boundary of the capacitor voltage across the flyover; yes k The voltage across the capacitor at time +1 This is the reference voltage value of the flying capacitor. It is the minimum flying capacitance value function.

[0024] As an alternative implementation, in the midpoint potential voltage balance layer, the designed value function for:

[0025] ;

[0026] in, yes k Midpoint potential voltage deviation at time +1.

[0027] As an alternative implementation method, the steady-state reward function for:

[0028] ;

[0029] Where L is the time window L; , and Current deviation value Cross capacitor voltage deviation value Midpoint potential voltage deviation value The weights; x represents the three phases a, b, and c.

[0030] As an alternative implementation method, the transient reward function for:

[0031] ;

[0032] in, h 1. h 2 and h 3 is the unit normalized value; r 4.r 5 and r 6 represents the weights for current tracking, flying capacitor voltage balancing, and midpoint potential voltage balancing; For reference current, This is the reference voltage value for the flying capacitor; and For a moment k The sampled values ​​of the current and the voltage across the capacitor; denoted as the midpoint potential voltage deviation; x represents the three phases a, b, and c.

[0033] Secondly, the present invention provides a multilevel inverter control system based on reinforcement learning sequential MPC, comprising:

[0034] The sequential result construction module is configured to construct a three-layer sequential optimization structure that progresses sequentially: current tracking, flying capacitor voltage balancing, and midpoint potential voltage balancing. Each layer solves for the optimal value based on its own value function and selects a candidate vector set according to the tolerance boundary. Each layer optimizes within the candidate vector set selected by the upper layer to generate the optimal control command.

[0035] The reinforcement learning model building module is configured to take current, flying capacitor voltage, midpoint potential voltage and corresponding deviation values ​​as observation inputs, and current tracking tolerance boundary and flying capacitor voltage balance tolerance boundary as action outputs. A multi-objective reward function combining steady-state reward function and transient reward function is designed.

[0036] The control module is configured to train a reinforcement learning agent offline using a multi-objective reward function. The trained reinforcement learning agent determines the optimal tolerance boundary values ​​of the current tracking layer and the fly-through capacitor balance layer based on real-time observations, thereby achieving dynamic balance between current tracking, fly-through capacitor voltage, and midpoint voltage.

[0037] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.

[0038] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.

[0039] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] This invention proposes a multi-level inverter control method and system based on reinforcement learning-based sequential MPC, simultaneously optimizing multiple control objectives such as current tracking, flying capacitor voltage balance, and midpoint potential voltage balance. Key objectives are optimized independently layer by layer, with each layer operating within the candidate vector set filtered by its predecessor. Interpretable tolerance boundary values ​​are used instead of weighting factors; these values ​​directly represent the allowable deviation boundary based on the optimal solution, avoiding objective conflicts inherent in traditional weighted summation. Simultaneously, based on real-time control objective deviations, the reinforcement learning agent dynamically adjusts its output tolerance boundary values, balancing response speed and stability, dynamically expanding the solution space, and breaking fixed-value constraints. The independent handling of equal physical quantities by each decision layer avoids the drawbacks of compromises between different dimensions. When the objective changes suddenly, each layer independently optimizes control. Furthermore, this invention's method, based on offline reinforcement learning's autonomous learning, achieves automatic design and optimization of multi-objective tolerance boundary values, solving the key problems of low efficiency and blind spots in manual parameter tuning in multi-objective systems.

[0042] This invention constructs a three-layer sequential optimization structure of current tracking → flying capacitor voltage balance → midpoint potential voltage balance, forming a progressive optimization decision chain. The optimization is performed independently at each level, and the candidate vector set is constrained layer by layer. This eliminates the target conflict problem of the traditional weighted summation method and can be extended to any multi-level topology.

[0043] This invention uses a reinforcement learning adaptive tolerance boundary value generation mechanism to analyze current tracking deviation and capacitor voltage deviation in real time, dynamically generate time-varying tolerance boundary values, break through the limitations of traditional fixed tolerance boundary values, and enable the vector selection range to adapt to changes in operating conditions.

[0044] This invention uses offline training to automatically design weighted factors. Compared with traditional model predictive control, which requires a lot of manual parameter tuning based on experience, this invention achieves autonomous learning capabilities and fundamentally solves the problem of blind spots in parameter tuning for multi-objective coupled systems.

[0045] This invention constructs a hybrid architecture of MPC execution and deep reinforcement learning decision-making, and improves the dynamic response capability of the system through sequential predictive control and dynamic tolerance boundary value construction.

[0046] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0048] Figure 1 This is a flowchart of the multilevel inverter control method based on reinforcement learning sequential MPC provided in Embodiment 1 of the present invention;

[0049] Figure 2 This is a structural diagram of a novel hybrid seven-level converter system cascaded with a three-level T-type converter and an H-bridge;

[0050] Figure 3 This is a schematic diagram of the overall structural control provided in Embodiment 1 of the present invention;

[0051] Figure 4 This is a simulation diagram of the three-phase output current provided in Embodiment 1 of the present invention;

[0052] Figure 5 This is a simulation diagram of the flying capacitor voltage provided in Embodiment 1 of the present invention;

[0053] Figure 6 This is a simulation diagram of the midpoint potential voltage provided in Embodiment 1 of the present invention;

[0054] Figure 7 The output line voltage simulation diagram provided in Embodiment 1 of the present invention. Detailed Implementation

[0055] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0056] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0057] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. Furthermore, it should be understood that the terms “comprising” and “including”, and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0058] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0059] In the field of power electronic converter control, multilevel inverters need to achieve multiple objectives, such as output current tracking and capacitor voltage balance, forming a multi-objective optimization problem with mutual coupling. Traditional model predictive control widely adopts weighted factor fusion of objectives, but there are differences in physical dimensions, often relying on trial and error for parameter tuning based on human experience, making it difficult to establish intuitive physical relationships between the values. This ambiguity weakens the transparency of the algorithm and is more likely to cause response hysteresis in multi-objective dynamic conflict scenarios. At the same time, the allocation of fixed weights can trigger dynamic conflicts between control objectives, which may cause output waveform distortion or capacitor voltage instability, making it difficult to adapt to a wide range of operating conditions.

[0060] To address this issue, this invention proposes a multi-level inverter control method based on reinforcement learning-based sequential MPC. This method abandons the traditional weighting mechanism and constructs a hierarchical decision-making architecture: output voltage tracking serves as the primary decision layer, flying capacitor voltage balance as the secondary decision layer, and midpoint potential voltage balance as the final decision layer, performing sequential selection. Interpretable tolerance boundary values ​​are used instead of weighting factors; these values ​​directly represent the allowable deviation boundary based on the optimal solution. The solution space is dynamically relaxed through the output of the reinforcement learning agent, breaking the fixed-value constraint. Each decision layer independently handles the same physical quantities, avoiding the defects of compromise between different dimensions. When the target changes suddenly, each layer independently optimizes control.

[0061] Example 1

[0062] This embodiment provides a multilevel inverter control method based on reinforcement learning sequential MPC, such as... Figure 1 As shown, it includes:

[0063] A three-layer sequential optimization structure is constructed, which progresses sequentially from current tracking to flying capacitor voltage balancing to midpoint potential voltage balancing. Each layer solves for the optimal value based on its own value function and selects a candidate vector set according to the tolerance boundary. Each layer optimizes within the candidate vector set selected by the upper layer to generate the optimal control command.

[0064] The observation inputs are current, flying capacitor voltage, midpoint potential voltage and corresponding deviation values. The action outputs are current tracking tolerance boundary and flying capacitor voltage balance tolerance boundary. A multi-objective reward function combining steady-state reward function and transient reward function is designed.

[0065] The reinforcement learning agent is trained offline using a multi-objective reward function. The trained reinforcement learning agent is then used to determine the optimal tolerance boundary values ​​for the current tracking layer and the flying capacitor balance layer based on real-time observations, thereby achieving dynamic balance between current tracking, flying capacitor voltage, and midpoint voltage.

[0066] In this embodiment, a hybrid seven-level converter (T2C-HB) system is used as the controlled object, such as... Figure 2As shown, each phase consists of a three-level T-type converter cascaded with an H-bridge, the DC side includes two identical series capacitors, and each phase has a flying capacitor.

[0067] To address the multi-objective control requirements of the T2C-HB inverter, a three-layer sequential optimization structure is constructed based on the priority of the control objectives: current tracking → flying capacitor (FC) voltage balance → neutral point (NP) voltage balance. Each optimization level solves for the optimal value based on its own value function and generates a new candidate vector set through tolerance boundaries, which serves as the candidate input for subsequent levels. Each level optimizes within the candidate vector set selected by the upper level, forming a progressive optimization decision chain that constrains the candidate vector set layer by layer, avoiding the objective conflict problem of the weighted summation method. At the same time, the optimization is performed in the three-phase space vector diagram, and only the expression between the reference output voltage and the switching function needs to be modified, which can be extended to any multilevel converter.

[0068] The process consists of three layers: The first layer is current tracking: calculating the predicted deviation between the reference current and the actual current to generate a second-layer candidate vector set that meets current accuracy requirements. The second layer is FC voltage balancing: selecting a third-layer candidate vector set from the second-layer candidate vector set to suppress voltage fluctuations across the flying capacitor. The third layer is NP voltage balancing: selecting the optimal vector from the third-layer candidate vector set of FC voltage balancing to maintain stable DC midpoint potential voltage, thereby obtaining the optimal control command.

[0069] Specifically:

[0070] In the first layer, to ensure current tracking, the designed value function is:

[0071] (1);

[0072] The constraints that the second-level candidate vector set must satisfy are:

[0073] (2);

[0074] in, This represents the maximum permissible deviation of the current value function, i.e., the current tracking tolerance boundary. and For reference current, and yes k The current at time +1 The minimum current value function is obtained by evaluating all candidate vectors. .

[0075] In this layer, a second layer of candidate vector sets is generated based on the current tracking tolerance boundary, within the allowable deviation range.

[0076] In the second layer, to ensure voltage balance of the flying capacitor FC, the designed value function is:

[0077] (3);

[0078] The constraints that the candidate vector set in the third layer must satisfy are:

[0079] (4);

[0080] in, The maximum permissible deviation of the FC voltage value function, i.e., the FC voltage balance tolerance boundary; yes k FC voltage at time +1 This is the FC reference voltage value. The minimum fly-through capacitance value function is obtained by evaluating all candidate vector sets in the second layer. .

[0081] In this layer, a third-layer candidate vector set is generated within the allowable deviation range based on the FC voltage balance tolerance boundary.

[0082] In the third layer, to ensure the balance of the midpoint potential NP voltage, the designed value function is:

[0083] (5);

[0084] in, yes k NP voltage deviation at time +1.

[0085] The objective function is minimized within the third-level candidate vector set, and the optimal control command is output. This achieves independent optimization of the objective at each level, avoiding the problems of numerous weight tuning experiments and dynamic performance degradation caused by the differences in physical dimensions of traditional multi-objective weight factors.

[0086] In this embodiment, the optimal tolerance boundary policy is learned using the Reinforcement Learning (RL) TD3 (Twin Delayed Deep Deterministic policy gradient algorithm). First, the system state is sampled in real time, including current, flyby capacitor voltage, midpoint potential voltage, and corresponding deviations. This data is used as the observation input to track the tolerance boundary with the current. and flying capacitor voltage balance tolerance boundary The action output is determined, and a function reward is set according to the desired control performance to minimize the total harmonic distortion (THD), balance the FC voltage, and balance the NP voltage.

[0087] Therefore, a steady-state reward function is designed. and transient reward function Combined multi-objective reward function R Guiding the direction of learning is defined as:

[0088] (6).

[0089] The steady-state reward function employs continuous offset detection and introduces a time window accumulator to accumulate the deviations of current, FC voltage, and NP voltage over L consecutive control cycles, defined as:

[0090] (7);

[0091] ;

[0092] in, , , and For a moment k The sampled values ​​of current, FC voltage, and NP voltage; The reference current is used; the time window L is set to 20, which serves as the time range for suppressing oscillations and enhancing the stability and performance of the entire control system. , and These are the deviation values ​​for current, FC voltage, and NP voltage, respectively. , and Current deviation value FC voltage deviation value NP voltage deviation value The weights; x represents the three phases a, b, and c.

[0093] The transient reward function is defined as a weighted sum of the deviation values ​​at the current time step:

[0094] (8);

[0095] in, h 1, h 2 and h 3 is the unit normalization value, which standardizes the error into a dimensionless value; r 4. r 5 and r 6 represents the weights for current tracking, FC voltage balance, and NP voltage balance, with a value range of [0,1]. These weights serve as the weight coefficients for the three control objectives, and their magnitudes vary with the modulation index. m Changes, when mWhen the current is relatively small, the main control task of the agent is to suppress current errors and reduce the total harmonic distortion (THD) of the current; conversely, when the current is relatively large... m When the voltage is large, the agent's main control task is to balance the FC voltage and reduce it. fluctuation.

[0096] This embodiment employs a design strategy that superimposes steady-state and transient reward functions. The steady-state reward function is constructed based on the integral of the error over a moving-time window and is used to characterize the system's continuous deviation state. The transient reward function directly captures the dynamic deviation at the current moment. When the current tracking deviation increases, or when the voltage across the flying capacitor or the midpoint potential experiences a continuous shift, the corresponding penalty term will correct the control strategy. The multi-objective reward function guides the control strategy to dynamically adjust towards multi-parameter collaborative optimization, thereby achieving autonomous learning and optimization of the system.

[0097] This enables the agent to adjust according to multiple control objectives, enhancing its adaptability and flexibility to different modulation variations. Simultaneously, a mechanism for generating modulation schemes based on random environments is designed to provide the reinforcement learning controller with ample training experience to cope with changing operating conditions.

[0098] The coefficients of each item are defined as follows:

[0099] (9);

[0100] in, r 4max Set it to 0.75 to ensure r 4 is greater than 0 to prevent the agent from making incorrect control actions during training, reflects the maximum weight allocation of the harmonic suppression task in the reward function, and indicates that the system gives the highest priority to the quality of the output current waveform. α This indicates the priority of the two control objectives: FC voltage balance and NP voltage balance; the adjustment coefficient. k It is a proportionality coefficient used to adjust the magnitude of change, and its optimal value is... kopt Experiments have determined that when k = kopt It can achieve Pareto optimal balance of current tracking, FC voltage balance, and NP midpoint balance.

[0101] like Figure 3The diagram illustrates the overall framework of the automatic tolerance boundary value design method. The TD3 algorithm employed is a deterministic deep reinforcement learning algorithm within the Actor-Critic (AC) framework, combining a deep deterministic policy gradient algorithm with dual Q-learning. Specifically, continuous sampling of current, flying capacitor voltage, midpoint potential voltage, and corresponding deviation values ​​is used as the agent's state variables. These state variables are mapped to multi-dimensional feature vectors and input into the policy and value networks of the reinforcement learning agent in real time. The multi-objective reward function is calculated in real time based on the current system state, and the optimal tolerance boundary value is output to the MPC controller through autonomous learning. After executing a control action, the agent enters a new state, forming a closed-loop interaction mechanism. Through continuous iteration, the RL agent gradually explores a dynamic optimization strategy. A platform was built using MATLAB / Simulink, and the agent was trained offline using the Reinforcement Learning Toolbox before being loaded into the experimental equipment.

[0102] Simulation results are as follows Figures 4-7 As shown, where, Figure 4 In the middle, the current exhibits a sinusoidal pattern, with the amplitude changing abruptly from 3A to 5A; Figure 5 In the middle, the voltage across the flying capacitor is 1 / 4 of that on the DC side, fluctuating between 0.2V and 0.4V. Figure 6 The midpoint voltage is half that of the DC side, fluctuating between 0.1V and 0.2V. Figure 7 The voltage level changed from five levels to nine levels, which is the line voltage, confirming that the Agent can adaptively and dynamically adjust the tolerance boundary value and key control parameters according to changes in operating conditions.

[0103] This embodiment constructs a three-layer sequential optimization structure of current tracking → flying capacitor voltage balance → midpoint potential voltage balance, forming a progressive optimization decision chain, which constrains the candidate vector set layer by layer, avoiding the target conflict problem of the weighted summation method, and can be extended to any multi-level topology.

[0104] A condition awareness mechanism is constructed to dynamically capture the inverter's operating conditions (low-profile / high-profile) based on the real-time monitoring module, providing environmental status input for the dual-delay deep deterministic strategy gradient TD3.

[0105] The design incorporates a multi-objective reward function that combines steady-state and transient metrics to guide the learning direction, penalizing deviations in current tracking, overshoot capacitor voltage, or sustained midpoint potential shifts. The TD3 strategy network outputs the optimal combination of tolerance boundary values ​​for the current layer and the FC layer, selecting tolerance boundary values ​​suitable for the current operating conditions and avoiding tedious parameter tuning. The strategy network simultaneously outputs two tolerance boundary values ​​for the current tracking layer and the overshoot capacitor balancing layer, achieving dynamic balance between current tracking, overshoot capacitor voltage, and midpoint voltage. Specifically, under low-modulation conditions, the output tolerance boundary value of the current layer is automatically reduced to prioritize optimizing the current waveform quality; under high-modulation conditions, the output tolerance boundary values ​​of the current layer and the FC layer are simultaneously increased to prioritize capacitor voltage balance by relaxing control precision.

[0106] Example 2

[0107] This embodiment provides a multilevel inverter control system based on reinforcement learning sequential MPC, including:

[0108] The sequential result construction module is configured to construct a three-layer sequential optimization structure that progresses sequentially: current tracking, flying capacitor voltage balancing, and midpoint potential voltage balancing. Each layer solves for the optimal value based on its own value function and selects a candidate vector set according to the tolerance boundary. Each layer optimizes within the candidate vector set selected by the upper layer to generate the optimal control command.

[0109] The reinforcement learning model building module is configured to take current, flying capacitor voltage, midpoint potential voltage and corresponding deviation values ​​as observation inputs, and current tracking tolerance boundary and flying capacitor voltage balance tolerance boundary as action outputs. A multi-objective reward function combining steady-state reward function and transient reward function is designed.

[0110] The control module is configured to train a reinforcement learning agent offline using a multi-objective reward function. The trained reinforcement learning agent determines the optimal tolerance boundary values ​​of the current tracking layer and the fly-through capacitor balance layer based on real-time observations, thereby achieving dynamic balance between current tracking, fly-through capacitor voltage, and midpoint voltage.

[0111] It should be noted that the above modules correspond to the steps described in Embodiment 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0112] In further embodiments, the following is also provided:

[0113] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in Embodiment 1. For brevity, further details are omitted here.

[0114] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0115] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.

[0116] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.

[0117] The method in Example 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.

[0118] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.

[0119] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.

[0120] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.

[0121] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.

[0122] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.

[0123] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A multi-level inverter control method based on reinforcement learning sequential MPC, characterized in that, include: A three-layer sequential optimization structure is constructed, which progresses sequentially from current tracking to flying capacitor voltage balancing to midpoint potential voltage balancing. Each layer solves for the optimal value based on its own value function and selects a candidate vector set according to the tolerance boundary. Each layer optimizes within the candidate vector set selected by the upper layer to generate the optimal control command. The observation inputs are current, flying capacitor voltage, midpoint potential voltage and corresponding deviation values. The action outputs are current tracking tolerance boundary and flying capacitor voltage balance tolerance boundary. A multi-objective reward function combining steady-state reward function and transient reward function is designed. The reinforcement learning agent is trained offline using a multi-objective reward function. The trained reinforcement learning agent is then used to determine the optimal tolerance boundary values ​​of the current tracking layer and the flying capacitor balance layer based on real-time observations, thereby achieving a dynamic balance between current tracking, flying capacitor voltage, and midpoint potential voltage. In the current tracking layer, a value function is designed based on the deviation between the reference current and the actual current. All candidate vectors are evaluated in this way, the minimum value function is determined, and a second set of candidate vectors within the allowable deviation range is generated based on the current tracking tolerance boundary. In the flying capacitor voltage balancing layer, a value function is designed based on the deviation between the reference flying capacitor voltage and the actual flying capacitor voltage to determine the minimum value function in the second-layer candidate vector set, and a third-layer candidate vector set within the allowable deviation range is generated based on the flying capacitor voltage balancing tolerance boundary. In the midpoint potential voltage balance layer, the optimal vector that maintains the DC midpoint potential voltage stability is selected from the candidate vector set in the third layer, thereby obtaining the optimal control command; In the current tracking layer, the value function of the design for: ; The constraints satisfied by the second-level candidate vector set are: ; in, For current tracking tolerance boundaries; and for k The reference current at time +1 and yes k The current at time +1 It is the minimum current value function; The value function of the design across the capacitor voltage balance layer. for: ; The constraints satisfied by the third-level candidate vector set are: ; in, To balance the tolerance boundary of the capacitor voltage across the flyover; yes k The voltage across the capacitor at time +1 This is the reference voltage value of the flying capacitor. This is the minimum flying capacitance value function.

2. The multi-level inverter control method based on reinforcement learning sequential MPC as described in claim 1, characterized in that, The value function designed at the midpoint potential-voltage balance layer. for: ; in, yes k Midpoint potential voltage deviation at time +1.

3. The multilevel inverter control method based on reinforcement learning sequential MPC as described in claim 1, characterized in that, Steady-state reward function for: ; Where L is the time window; , and The current deviation value at time k Cross capacitor voltage deviation value and midpoint potential voltage deviation value The weights; x represents the three phases a, b, and c.

4. The multi-level inverter control method based on reinforcement learning sequential MPC as described in claim 1, characterized in that, Transient reward function for: ; in, h 1. h 2 and h 3 is the unit normalized value; r 4. r 5 and r 6 represents the weights for current tracking, flying capacitor voltage balancing, and midpoint potential voltage balancing; Let be the reference current at time k. This is the reference voltage value for the flying capacitor; and These are the sampled values ​​of the current and the voltage across the flying capacitor at time k; Let be the midpoint potential voltage deviation value at time k; x represents the three phases a, b, and c.

5. A multilevel inverter control system based on reinforcement learning sequential MPC, characterized in that, include: The sequential result construction module is configured to construct a three-layer sequential optimization structure that progresses sequentially: current tracking, flying capacitor voltage balancing, and midpoint potential voltage balancing. Each layer solves for the optimal value based on its own value function and selects a candidate vector set according to the tolerance boundary. Each layer optimizes within the candidate vector set selected by the upper layer to generate the optimal control command. In the current tracking layer, a value function is designed based on the deviation between the reference current and the actual current. All candidate vectors are evaluated in this way, the minimum value function is determined, and a second set of candidate vectors within the allowable deviation range is generated based on the current tracking tolerance boundary. In the flying capacitor voltage balancing layer, a value function is designed based on the deviation between the reference flying capacitor voltage and the actual flying capacitor voltage to determine the minimum value function in the second-layer candidate vector set, and a third-layer candidate vector set within the allowable deviation range is generated based on the flying capacitor voltage balancing tolerance boundary. In the midpoint potential voltage balance layer, the optimal vector that maintains the DC midpoint potential voltage stability is selected from the candidate vector set in the third layer, thereby obtaining the optimal control command; In the current tracking layer, the value function of the design for: ; The constraints satisfied by the second-level candidate vector set are: ; in, For current tracking tolerance boundaries; and for k The reference current at time +1 and yes k The current at time +1 It is the minimum current value function; The value function of the design across the capacitor voltage balance layer. for: ; The constraints satisfied by the third-level candidate vector set are: ; in, To balance the tolerance boundary of the capacitor voltage across the flyover; yes k The voltage across the capacitor at time +1 This is the reference voltage value of the flying capacitor. The minimum cross-capacitance value function; The reinforcement learning model building module is configured to take current, flying capacitor voltage, midpoint potential voltage and corresponding deviation values ​​as observation inputs, and current tracking tolerance boundary and flying capacitor voltage balance tolerance boundary as action outputs. A multi-objective reward function combining steady-state reward function and transient reward function is designed. The control module is configured to train a reinforcement learning agent offline using a multi-objective reward function. The trained reinforcement learning agent determines the optimal tolerance boundary values ​​of the current tracking layer and the fly-through capacitor balance layer based on real-time observations, thereby achieving dynamic balance between current tracking, fly-through capacitor voltage, and midpoint potential voltage.

6. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-4.

8. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-4.