Power loss optimization control method and system for modular multilevel converter

By applying the Q-learning algorithm of the reinforcement learning model in MMC and dynamically optimizing the sorting frequency and balance adjustment number, the problems of limited loss balancing effect and system oscillation in traditional control methods are solved, and optimal power loss control and stable operation of MMC under different environments are achieved.

CN120675423APending Publication Date: 2025-09-19JIANGSU VOCATIONAL & TECHNICAL UNIVERSITY OF ARCHITECTURE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510817946.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional MMC power loss control methods cannot simultaneously optimize the coupling relationship between the sorting frequency and the balance adjustment number, resulting in limited loss balancing effects. In addition, when sub-modules fail or multiple operating conditions are switched, the PI parameters need to be frequently calibrated, causing system oscillations.

Method used

The Q-learning algorithm of the reinforcement learning model is used to select actions from the action space. The sorting frequency and balance adjustment number of MMC are optimized through multiple rounds of iterations. A closed-loop training mechanism of state-action-reward is constructed to maximize the global cumulative reward and achieve the optimal control strategy.

Benefits of technology

It achieves optimal power loss control of MMC under different operating environments, autonomously adapts to sub-module failures and operating condition switching, avoids the oscillation risk caused by frequent PI parameter calibration, and significantly enhances operational stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675423A_ABST
    Figure CN120675423A_ABST
Patent Text Reader

Abstract

The invention discloses a power loss optimization control method and system for a modular multilevel converter, and relates to the technical field of modular multilevel converters. Comprising the following steps: acquiring an initial state, and presetting an action space composed of various actions; based on the initial state, selecting an action from the action space through a Q learning algorithm of a reinforcement learning model, and determining a reward; the final states of the modular multilevel converter in different operation environments are obtained through multi-round iteration; taking the sum of the rewards of the maximized modular multilevel converter in different operation environments as an optimization target to train the reinforcement learning model to obtain a power loss optimization control model; and inputting the real-time state of the modular multilevel converter in the current operation environment into the power loss optimization control model to obtain an optimal power loss optimization control method. According to the invention, flexible adjustment can be realized according to real-time fault conditions and system states, and power loss balance is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of modular multi-level converters, and in particular to a method and system for optimizing power loss control of a modular multi-level converter. Background Art

[0002] In modular multilevel converters (MMCs), loss imbalance can cause some components to operate at high loads for a long time, increasing overall power loss. f s and balance adjustments N ban , which can optimize switching frequency and loss distribution, reduce ineffective energy consumption, and improve system energy efficiency. Power loss optimization control is the core technology to ensure efficient and reliable operation of MMC.

[0003] The traditional MMC power loss control method uses a single degree of freedom regulation based on a PI controller. The principle of single degree of freedom regulation based on a PI controller is to use a closed-loop proportional-integral controller to adjust a single parameter (such as the sorting frequency). f s or balance adjustment N ban The control process is as follows: first, the bridge arm current, submodule capacitor voltage, power loss and other parameters are obtained through sensors. Then the target power loss is preset. P ref Or voltage balance threshold. , the PI controller outputs the regulation value (such as f s Increment Δf s ): Finally, the adjusted f s or N ban Applied to MMC control system.

[0004] The limitation of traditional MMC power loss control method is that it cannot optimize f s and N ban The coupling relationship between the two leads to limited loss balancing effect, which requires frequent calibration of PI parameters when a submodule fails or multiple working conditions are switched, thus causing system oscillation. Summary of the Invention

[0005] Based on this, it is necessary to provide a power loss optimization control method and system for a modular multi-level converter to address the above technical problems.

[0006] An embodiment of the present invention provides a method for optimizing power loss control of a modular multilevel converter, comprising: Acquiring an initial state of the modular multilevel converter in any operating environment, the initial state including a sorting frequency and a balancing adjustment number; and determining an action space using an adjustment strategy of the modular multilevel converter as an action; Based on the initial state, an action is selected from the action space using the Q-learning algorithm of the reinforcement learning model. The intermediate state of the modular multilevel converter in the next operating environment is determined based on the selected action. The reward is determined based on the minimum power loss error and the minimum power loss between the initial state and the intermediate state. Based on the intermediate state of the modular multilevel converter in the next operating environment, an action is reselected from the action space. Through multiple rounds of iteration, the final state of the modular multilevel converter in different operating environments is obtained. The optimization objective is to maximize the sum of the rewards of the modular multilevel converter in different operating environments to train the reinforcement learning model and obtain a power loss optimization control model. The real-time status of the modular multilevel converter in the current operating environment is input into the power loss optimization control model to obtain an optimal power loss optimization control method including an adjustment strategy of the modular multilevel converter.

[0007] Optionally, the initial state of the modular multilevel converter in any operating environment also includes: DC side voltage, grid voltage, power loss and proportion of faulty submodules The proportion of faulty submodules is determined based on the following formula: ; in, is the number of faulty submodules, N is the total number of submodules, is the proportion of faulty submodules.

[0008] Optionally, determining an action space using an adjustment strategy of a modular multilevel converter as an action specifically includes: According to the values ​​of the sorting frequency and the balance adjustment number quantization state, the value space of the sorting frequency and the value space of the balance adjustment number are determined based on the following formula: ; ; The values ​​of the sorting frequency and the balancing adjustment number are used as the adjustment strategy of the modular multilevel converter, and the action space is constructed based on the following formula: ; in, δ for s The quantized amount of sFor status, A is the action space, is the value space of sorting frequency, To adjust the value space of the balance number.

[0009] Optionally, the reward is determined based on the minimum power loss error and the minimum power loss between the initial state and the intermediate state by the following formula: ; in, r ( s , a ) as a reward, F is the objective function of minimum power loss error and minimum power loss, F old is the previous objective function value, F new is the current objective function value, r max is the maximum positive reward given when energy loss is reduced, r min The bonus given when the energy loss remains constant, r p Negative rewards given when energy loss increases.

[0010] Optionally, an action is selected from the action space by a Q-learning algorithm of a reinforcement learning model based on the following formula, and an intermediate state of the modular multilevel converter in the next operating environment is determined according to the selected action: ; in, Q k ( s , a ) is the state and the action selected according to the state a Next Q value, Q k ( s ', a ') is the next state and the action selected according to the next state Q value, s For status, a is the action selected based on the state, s 'For the next state, a ' is the action selected according to the next state, α is the learning rate, γ is the discount factor, Q k+1 ( s , a ) is the intermediate state of the modular multilevel converter in the next operating environment.

[0011] An embodiment of the present invention further provides a power loss optimization control system for a modular multilevel converter, comprising: a three-phase MMC main circuit, a control circuit, and a reinforcement learning module; Each phase of the three-phase MMC main circuit includes an upper bridge arm and a lower bridge arm, and each upper bridge arm and each lower bridge arm includes multiple submodules and bridge arm inductors connected in series; The control circuit includes a digital signal processor, a drive circuit, and a sensor module. The digital signal processor is connected to the drive circuit via optical fiber to control each submodule. The sensor module collects DC side voltage and grid voltage in real time, determines power loss, fault submodule ratio, sorting frequency, and balance adjustment number based on the DC side voltage and grid voltage, and feeds these back to the digital signal processor. The reinforcement learning module is embedded in the digital signal processor and is used to implement the power loss optimization control method of the modular multi-level converter.

[0012] The above-mentioned method and system for optimizing power loss control of a modular multi-level converter provided by the embodiments of the present invention have the following beneficial effects compared with the prior art: The present invention constructs a dual-parameter state space including sorting frequency and balance adjustment number, and constructs an action space with the adjustment strategy of modular multi-level converter as the action, uses Q learning algorithm to realize action selection, and automatically learns the optimal combination of sorting frequency and balance adjustment number in the iterative process, while optimizing f s and N ban The coupling relationship between the two modules breaks through the limitations of traditional MMC power loss control methods. Through multi-environment iterative training, the power loss optimization control model can autonomously adapt to sub-module failures and operating condition switching, avoiding the oscillation risk caused by frequent PI parameter calibration. In addition, the present invention adopts a state-action-reward closed-loop training mechanism with the goal of maximizing the global cumulative reward to obtain an optimal control strategy that takes into account both transient response and steady-state efficiency. It has the ability to independently handle complex constraints such as capacitor voltage fluctuations and circulating current suppression, significantly enhancing the operational stability of modular multilevel converters across the full range of operating conditions, and providing a smarter and more efficient loss control solution for flexible DC transmission systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A schematic diagram of a three-phase MMC simulation system for a power loss optimization control method of a modular multilevel converter provided in one embodiment; Figure 2 A three-phase MMC experimental circuit configuration diagram of a power loss optimization control method for a modular multilevel converter provided in one embodiment; Figure 3A zero-current schematic diagram of a power loss optimization control method for a modular multi-level converter provided in one embodiment; Figure 4 An MMC topology diagram of a power loss optimization control method for a modular multilevel converter provided in one embodiment; Figure 5 A schematic diagram of an iterative process between an intelligent agent and an environment in a power loss optimization control method for a modular multi-level converter provided in one embodiment; Figure 6 FIG. 1 is a diagram of an MMC loss balancing control strategy of a power loss optimization control method for a modular multilevel converter provided in one embodiment. DETAILED DESCRIPTION

[0014] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0015] Currently, power loss optimization has achieved some success, but the following issues remain: Hardware-based approaches increase system hardware cost and complexity. The loss adjustment range of reference signal-based adjustment methods is limited by the adjustment range of the modular multilevel converter. This effectiveness is limited when the adjustment range is large. Methods based on submodule capacitor voltage control can increase submodule capacitor voltage ripple. Current power loss optimization only considers loss optimization during normal MMC operation, ignoring loss optimization in asymmetric bridge arm structures.

[0016] While power loss optimization control addresses loss imbalance at the submodule and device levels during normal operation, it addresses the arm-level loss imbalance when dealing with asymmetric bridge arm structures, making standard loss optimization methods unsuitable for this specific scenario.

[0017] Traditional MMC power loss control methods utilize single-degree-of-freedom regulation based on a PI controller. This limitation lies in the ability to adjust parameters within a single degree of freedom. Conventional loss balance control methods often utilize PI controllers, which inherently present risks of parameter calibration and system instability. When sub-module (SM) failures are frequent, greater degrees of freedom need to be adjusted, increasing the potential for system instability. Furthermore, traditional MMC power loss control methods require the establishment of a linear relationship between loss and parameters, making them difficult to handle for nonlinear or time-varying systems.

[0018] In one embodiment, a method for optimizing power loss control of a modular multilevel converter is provided, the method comprising: An initial state of the modular multilevel converter in any operating environment is obtained, wherein the initial state includes a sorting frequency and a balance adjustment number, and an action space is determined with an adjustment strategy of the modular multilevel converter as an action.

[0019] Based on the initial state, the Q-learning algorithm (Q-1 learning) of the reinforcement learning model selects an action from the action space. The selected action determines the intermediate state of the modular multilevel converter in the next operating environment. The reward is determined based on the minimum power loss error and the minimum power loss between the initial state and the intermediate state.

[0020] Based on the intermediate state of the modular multilevel converter in the next operating environment, actions are reselected from the action space. Through multiple rounds of iteration, the final state of the modular multilevel converter in different operating environments is obtained. Maximizing the sum of the rewards of the modular multilevel converter in different operating environments is used as the optimization objective to train the reinforcement learning model and obtain a power loss optimization control model.

[0021] The real-time status of the modular multilevel converter in the current operating environment is input into the power loss optimization control model to obtain an optimal power loss optimization control method including an adjustment strategy of the modular multilevel converter.

[0022] The specific processing process includes: 1. Data Collection and Preprocessing The rated voltage of SM during normal operation is expressed as: in, V dc is the DC side voltage, N is the total number of submodules.

[0023] Assume that a fault occurs in the upper arm submodule of phase A of the MMC. The faulty SM is bypassed and a redundant SM is placed in the bridge arm to maintain the normal operation of the MMC when the SM fails. The proportion of faulty SMs is:

[0024] ; in, is the number of faulty submodules, N is the total number of submodules.

[0025] Therefore, the voltage of the SM capacitor in the fault bridge arm is: ; According to the above formula, when SM fault occurs, the SM capacitor voltage U f Changes and γ When γ When increasing, Uf Increase, when γ When reducing, U f reduce.

[0026] According to the voltage balance control strategy of BAN, the switching frequency of each SM in the bridge arm is calculated. f sw : ; in m is the modulation coefficient, T f is the fundamental frequency, f s is the sort frequency.

[0027] When a SM fault occurs, the SM switching frequency f sw Changes and f s and N ban When fs / Nban When increasing, f sw Increase, when f s / N ban When reducing, f sw Decrease.

[0028] When MMC adopts fundamental frequency circulating current suppression and dual frequency circulating current suppression, the A phase upper arm current within one fundamental wave cycle can be obtained. i au .like Figure 3 As shown, i au The zero current crossing point t 1. t 2 and t 3 can be expressed as:

[0029] ; SM medium power devices T 1. T 2. D 1 and D The absolute average value of the current of 2 is: ; The power loss of MMC is mainly caused by the switching T 1 / T 2 and diode D 1 / D 2, including conduction loss and switching loss. Four power devicesT 1. T 2. D 1 and D The conduction loss of 2 can be expressed as:

[0030] ; in Vce0 and Vd0 They are the zero current voltage drop of IGBT and diode in the on state respectively. Rce0 and Rd0 are the on-resistances of the IGBT and diode respectively.

[0031] According to the above analysis, the conduction loss is almost unaffected by the SM fault. The switching loss is caused by the average current and switching frequency of the SM semiconductor device, and its variation trend is similar to f sw Consistency .

[0032] T 1. T 2. D 1 and D 2 Total Power loss P T1 、 P D1 、 P T2 and P D2 It can be expressed as: ; in i T1ave 、 i T2ave 、 i D1ave and i D2ave They are T 1. T 2. D 1 and D 2 average current value.

[0033] The total loss of an SM is the sum of the conduction loss and switching loss of each device. Since an SM fault does not affect the conduction loss, the switching loss in the faulted leg increases. Therefore, the total loss increases when an SM fault occurs.

[0034] Based on the above analysis, the power loss P T1 / 2 and P D1 / 2 Can be adjusted f s andN ban to adjust.

[0035] Based on the above analysis, a power loss optimization control method for modular multilevel converter is proposed to balance the power loss of the asymmetric bridge arm of MMC caused by submodule failure. The loss value of the power device with the highest power loss in the submodule is taken as the power loss of the entire submodule. PSM=Max { P T1 ,P T2 ,P D1 ,P D2}, where power loss P T1 、 P T2 、 P D1 and P D2 The power loss of SMs in the healthy arm is symmetrical, and the reference value of power loss can be set to the power loss of any SM in the healthy arm. P l_ref .

[0036] Taking the upper arm of phase A as an example, the reference value of power loss in the loss balance control strategy can be set as ; in P l_jk ( j=a,b,c ; k=u,l ) is the power loss of the bridge arm.

[0037] Q -learning is an algorithm for machine learning that can handle random transitions and rewards without the need for an environment model and tuning. Q -Learning algorithms generally include five concepts: agent, environment, state, action and reward. s ,action a and rewards r Autonomously interact with the environment. Rewards can be seen as feedback reinforcement signals that reflect the quality of the choices made. As the agent continues to interact with the environment, Q -Learning can accumulate experience through feedback signals and continuously learn how to choose the best actions. The goal of the agent is to converge to the most appropriate action strategy to achieve the maximum total reward. Figure 5 The following is the iterative process between the agent and the environment in Q-learning. Q-In learning algorithms, experience is stored in Q This helps determine the best strategy. Q The value table consists of the transition probabilities of different states, and the algorithm selects Q The action with the highest value is used to maximize performance. Therefore, Q -learning algorithms are well suited to rapidly determine the optimal control variables for MMC in this study ( fs and Nban ), aimed at balancing the MMC losses throughout its operating range. Q -learning, constructing state space and action space, defining rewards and Q The value update formula helps to quickly and accurately obtain the optimal control variables of MMC fs and Nban .

[0038] exist Q -In the learning algorithm, there are three elements, namely state, action and reward. MMC In the initial state, the power loss P l_jk , Faulty submodule ratio , sorting frequency f s and balance adjustments N ban composition.

[0039] 2. State space definition Power loss under current input conditions P l_jk ( j=a,b,c ; k=u,l ) by the current f s and N ban Therefore, the state space is defined as S :

[0040] ; in: f s : Sorting frequency, which determines the update period of submodule voltage sorting (unit: Hz) N ban : Balance adjustment number, the number of submodules that need to be adjusted in each sorting cycle 3. Action Space and Strategy Selection state s The transition is determined by the current action a The adjustment strategy of modular multilevel converterπ , from the current state s Get the intermediate state. In this process, s The value of f s and N ban Therefore, by adjusting f s and N ban The value of the next state s '. In addition, it should be based on f s 、 N ban sensitivity to quantify s The value of f s 、 N ban The value space is defined as:

[0041] ; ; in δ yes s The quantized amount of f s and N ban The increment Δ( f s / N ban ) should satisfy the constraints: ; Therefore, the action space is defined as: ; 4. Reward Design The reward is determined based on the minimum power loss error and the minimum power loss between the initial state and the intermediate state: ; in, r ( s , a ) as a reward, F is the objective function of minimum power loss error and minimum power loss, F old is the previous objective function value, F new is the current objective function value, r max is the maximum positive reward given when energy loss is reduced, r minis the reward given when the energy loss remains constant, r p is a negative reward given when energy loss increases.

[0042] pass F Evaluate MMC performance, F The smaller the value, the better the loss leveling control performance.

[0043] ; in P A ( f s , N ban ) is the power loss function, ΔP ( f s , N ban ) is the power loss error function, φ Is the penalty factor. Since it is difficult to directly use nonlinear equality constraints in the Q-1earning algorithm P0'-P0 , so the power loss error function is defined as ΔP :

[0044] ; in represents the power loss during training, P 0 is the expected power loss.

[0045] 5. Q Value update and strategy optimization The Q-learning algorithm of the reinforcement learning model selects an action from the action space, and the intermediate state of the modular multilevel converter in the next operating environment is determined based on the selected action: ; in Q k ( s , a ) is the state and the action selected according to the state a Next Q value, Q k ( s ', a ') is the next state and the action selected according to the next state Q value, s For status, a is the action selected based on the state, s 'For the next state, a ' is the action selected according to the next state,α is the learning rate, γ is the discount factor, Q k+1 ( s , a ) is the intermediate state of the modular multilevel converter in the next operating environment.

[0046] Strategy convergence: Based on the intermediate state of the modular multilevel converter in the next operating environment, the action is reselected from the action space. Through multiple rounds of iteration, the final state of the modular multilevel converter in different operating environments is obtained. Maximizing the sum of the rewards of the modular multilevel converter in different operating environments is used as the optimization objective to train the reinforcement learning model and obtain a power loss optimization control model.

[0047] Through multiple iterations of training, Q The value table gradually converges to the optimal strategy and generates control parameters f s and N ban adjustment instructions.

[0048] Table 1 Q- Key parameters of the Learing algorithm In order to achieve MMC The optimal loss balancing effect of the present invention is achieved by ε -greedy method selects behavior, explores more strategies, and saves the optimal power loss state in each strategy. ε -After learning the greedy method, select the minimum objective function value obtained by N rounds of training F min As a parameter for state update. Then, use Q The behavior selection method is used to continue training until the policy converges.

[0049] After the Q-learning algorithm training process is completed, the training results are stored in a lookup table. The input of the lookup table is the current operating environment (initial state), including the sorting frequency f s , balance adjustment N ban and power loss P l_jk The output in the lookup table is the action policy response in the current state. In addition, f s and N banThe intervals are set to 10Hz and 1 respectively to reasonably control the accuracy and the size of the lookup table. In fact, when detecting the operating environment, it will first be quantified, and then the matching action strategy will be found directly from the lookup table. If the quantized operating environment cannot be found in this lookup table, the value closest to the current operating environment will be selected to directly find the corresponding action strategy. The above process is as follows Figure 6 As shown, it is based on Q -Learning MMC loss balance control strategy.

[0050] Based on the same inventive concept, a power loss optimization control system for a modular multilevel converter is also provided. The system includes a three-phase MMC main circuit, a control circuit, a reinforcement learning module and an experimental verification platform.

[0051] Each phase of the three-phase MMC main circuit includes an upper bridge arm and a lower bridge arm, and each upper bridge arm and each lower bridge arm includes a series N a The submodule adopts a half-bridge structure and integrates IGBT ( T 1, T 2) Anti-parallel diode ( D 1, D 2) and DC support capacitor C The topology of MMC is as follows: Figure 4 shown.

[0052] The control circuit includes a digital signal processor (DSP), a drive circuit, and a sensor module. The DSP is connected to the drive circuit via optical fiber to control the submodule switches. The sensor module collects DC side voltage and grid voltage in real time, determines power loss, fault submodule ratio, sorting frequency, and balance adjustment number based on the DC side voltage and grid voltage, and feeds back to the DSP.

[0053] The reinforcement learning module is embedded in the DSP, including the state space s =[ f s , N ban ], action space and Q value table are used to implement the power loss optimization control method of modular multilevel converter.

[0054] The experimental verification platform includes a MATLAB simulation model and a low-power MMC prototype. The prototype is built with FF75R12YT3 IGBT modules, and the DSP communicates with the host computer through the CAN bus. The system uses capacitor voltage balance control strategy and Q -Learning algorithm optimization significantly reduces the loss difference under sub-module failure.

[0055] Example 1: Simulation Platform Verification of Loss Leveling Control Under 10% Submodule Failures To verify the Q -learning algorithm based loss balance control strategy, and built a three-phase MMC using the professional tool MATLAB, such as Figure 1 The following is a schematic diagram of the simulation system. Set 5 submodules (SM46-SM50) in the upper bridge arm of phase A to fail (fault ratio γ =10%), triggering redundant module switching. Enable RL-PLBC algorithm based on Q-learning to dynamically adjust the sorting frequency f s and balance adjustments N ban . Initial sort frequency f s0 =2.4kHz, balance adjustment number N ban =10. After the algorithm is optimized by Q-learning, the initial sorting frequency f s0 down to 2.15kHz, N ban The power loss difference between the faulty bridge arm and the healthy bridge arm dropped from 0.418 kW to nearly 0 kW. The algorithm completed parameter optimization in 0.022 seconds, significantly reducing the fluctuations in the bridge arm current and capacitor voltage. The switching frequency of the faulty bridge arm dropped from 574 Hz to 470 Hz, a reduction of 18.5%, reducing thermal stress.

[0056] Example 2: Verifying submodule fault control under power surge using a simulation platform The system power suddenly drops from 1.0 pu to 0.8 pu, triggering a 10% submodule failure in the upper bridge arm of phase A. After the power sudden change, the RL-PLBC algorithm is enabled to jointly adjust the sorting frequency. f s and balance adjustments N ban . Initial sort frequency f s0 =2.4kHz, initial sorting frequency after optimization f s0 Increase to 2.31kHz; balance adjustment number N ban Adjusted from 10 to 9. Power loss equalization was achieved within 0.023 seconds, reducing the loss difference from 0.318kW to a balanced state. The submodule capacitor voltage fluctuation range was reduced from 3.8kV to 3.2kV, improving system stability. Furthermore, under the dual disturbances of power surges and faults, the algorithm was able to maintain the bridge arm current harmonic content below 2%.

[0057] Example 3: Experimental Platform Verification (Normal and Fault Conditions) The experimental platform includes a DSP controller, a DC power supply, and a single-phase MMC. A low-power MMC prototype is built using FF75R12YT3 IGBT modules, with each bridge arm containing 10 submodules (including 2 redundant ones). Figure 2 This is a three-phase MMC experimental circuit configuration. The DSPTMS320F28335 implements the RL-PLBC algorithm and communicates with the host computer via the CAN bus. An oscilloscope records the bridge arm current, capacitor voltage, and power loss waveforms in real time. System symmetry is verified when there is no fault, and then a submodule fault is manually triggered in the upper bridge arm of phase A ( γ =10%), RL-PLBC is enabled. The implementation effect is: the difference in power loss between the upper and lower bridge arms is less than 0.05kW under normal operating conditions, verifying the control baseline performance. Within 0.3 seconds after the fault, the initial sorting frequency f s0 Adjust from 2.4kHz to 2.15kHz, balance adjustment number N ban Adjusted from 10 to 9. The loss of the faulty bridge arm dropped from 3.11kW to 2.69kW, consistent with the healthy bridge arm, and the rise in IGBT junction temperature was reduced by 15%, extending the device life.

[0058] The above-described embodiments merely illustrate several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.

Claims

1. A method for optimizing power loss control of a modular multilevel converter, characterized in that: include: Acquiring an initial state of the modular multilevel converter in any operating environment, wherein the initial state includes a sorting frequency and a balance adjustment number; and determining the action space with the adjustment strategy of the modular multilevel converter as the action; Based on the initial state, an action is selected from the action space using the Q-learning algorithm of the reinforcement learning model. The intermediate state of the modular multilevel converter in the next operating environment is determined based on the selected action. The reward is determined based on the minimum power loss error and the minimum power loss between the initial state and the intermediate state. Reselect actions from the action space based on the intermediate state of the modular multilevel converter in the next operating environment, and obtain the final state of the modular multilevel converter in different operating environments through multiple rounds of iterations; Maximizing the sum of rewards for the modular multilevel converter in different operating environments is used as the optimization objective to train the reinforcement learning model and obtain a power loss optimization control model. The real-time status of the modular multilevel converter in the current operating environment is input into the power loss optimization control model to obtain an optimal power loss optimization control method including an adjustment strategy of the modular multilevel converter.

2. The method for optimizing power loss control of a modular multilevel converter according to claim 1, wherein: The initial state of the modular multilevel converter in any operating environment also includes: DC side voltage, grid voltage, power loss and proportion of faulty submodules; The proportion of faulty submodules is determined based on the following formula: ; in, is the number of faulty submodules, N is the total number of submodules, is the proportion of faulty submodules.

3. The power loss optimization control method of a modular multilevel converter according to claim 1, wherein: The determining of the action space using the adjustment strategy of the modular multilevel converter as an action specifically includes: According to the values ​​of the sorting frequency and the balance adjustment number quantization state, the value space of the sorting frequency and the value space of the balance adjustment number are determined based on the following formula: ; ; The values ​​of the sorting frequency and the balancing adjustment number are used as the adjustment strategy of the modular multilevel converter, and the action space is constructed based on the following formula: ; in, δ for s The quantized amount of s For status, A is the action space, is the value space of sorting frequency, To adjust the value space of the balance number.

4. The method for optimizing power loss control of a modular multilevel converter according to claim 1, wherein: The reward is determined based on the minimum power loss error and the minimum power loss between the initial state and the intermediate state using the following formula: ; in, r ( s , a ) as a reward, F is the objective function of minimum power loss error and minimum power loss, F old is the previous objective function value, F new is the current objective function value, r max is the maximum positive reward given when energy loss is reduced, r min The bonus given when the energy loss remains constant, r p Negative rewards given when energy loss increases.

5. The power loss optimization control method of a modular multilevel converter according to claim 4, characterized in that: Based on the following formula, an action is selected from the action space by the Q-learning algorithm of the reinforcement learning model, and the intermediate state of the modular multilevel converter in the next operating environment is determined according to the selected action: ; in, Q k ( s , a ) is the state and the action selected according to the state a Next Q value, Q k ( s ', a ') is the next state and the action selected according to the next state Q value, s For status, a is the action selected based on the state, s 'For the next state, a ' is the action selected according to the next state, α is the learning rate, γ is the discount factor, Q k+1 ( s , a ) is the intermediate state of the modular multilevel converter in the next operating environment.

6. A modular multilevel converter power loss optimization control system based on the modular multilevel converter power loss optimization control method according to any one of claims 1 to 5, characterized in that: include: Three-phase MMC main circuit, control circuit and reinforcement learning module; Each phase of the three-phase MMC main circuit includes an upper bridge arm and a lower bridge arm, and each upper bridge arm and each lower bridge arm includes multiple submodules and bridge arm inductors connected in series; The control circuit includes a digital signal processor, a drive circuit, and a sensor module. The digital signal processor is connected to the drive circuit via optical fiber to control each submodule. The sensor module collects DC side voltage and grid voltage in real time, determines power loss, fault submodule ratio, sorting frequency, and balance adjustment number based on the DC side voltage and grid voltage, and feeds these back to the digital signal processor. The reinforcement learning module is embedded in the digital signal processor and is used to implement the power loss optimization control method of the modular multi-level converter.