Pre-charging control method for high-voltage high-power suspension capacitor
By adopting a framework of fusion of deep learning and attention mechanisms in suspended capacitor pre-charge control, a pre-charge control model is built, which solves the problem that traditional methods are difficult to capture dynamic changes, and achieves high-precision voltage control and system stability, which is suitable for a variety of complex working conditions.
Patent Information
- Application Number
- CN202510550537.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Traditional suspended capacitor precharge control methods are difficult to capture the dynamic changes of voltage, current and temperature, resulting in the inability to quickly adjust the charging strategy when the bus voltage fluctuates or the temperature rises sharply, affecting the stability of the system, and it is difficult to achieve high-precision control in new scenarios, making it prone to overcharge or undercharge.
A pre-charge control model is constructed using a framework based on the fusion of deep learning and attention mechanisms. The feature vector is generated through the real-time acquisition of charging parameters, the Actor network is used to generate actions and long-term profit evaluation, the charging process is optimized, and the training is combined with meta-learning is carried out to adapt to multiple scenarios.
The voltage control accuracy is improved, ensuring stable operation of the system under complex operating conditions such as sudden temperature changes and bus voltage fluctuations, and reducing the occurrence of overcharge or undercharge.
Smart Images

Figure CN120074255A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of power electronics technology, and in particular, to a pre-charge control method for a high-voltage and high-power floating capacitor. Background Art
[0002] A floating capacitor refers to a capacitor used to improve the electric field distribution and balance the voltage in power electronic devices, especially in multilevel inverters. Currently, the control methods for pre-charging floating capacitors mainly adopt fuzzy control or PID control algorithms; however, traditional algorithms lack the ability to model long-term time series dependencies and are difficult to capture the dynamic change laws of voltage, current, and temperature. For example, when the bus voltage fluctuates or the temperature rises suddenly, traditional algorithms cannot quickly adjust the charging strategy, affecting the system stability. Moreover, existing algorithms need to be retrained with a large amount of data in new scenarios (such as sudden temperature changes and load fluctuations), cannot quickly adapt to dynamic changes, and are difficult to achieve high-precision control, and overcharging or undercharging phenomena are likely to occur during the charging process. Summary of the Invention
[0003] One of the purposes of the present application is to provide a pre-charge control method for a high-voltage and high-power floating capacitor that can solve at least one of the defects in the above background art.
[0004] To achieve at least one of the above purposes, the technical solution adopted by the present application is: a pre-charge control method for a high-voltage and high-power floating capacitor, which is applied to a resonant soft-switching push-pull topology circuit, and includes the following control steps: constructing a pre-charge control model based on a fusion framework of deep learning and attention mechanism; using the real-time collected charging parameters of the floating capacitor as the input of the pre-charge control model to generate a feature vector; based on the obtained feature vector, performing a long-term reward evaluation on the actions generated by the Actor network for adjusting the pre-charge process; and optimizing the long-term reward according to the real-time charging effect of the floating capacitor and outputting the optimized actions.
[0005] Preferably, the generation of the feature vector includes the following process: preprocessing the input charging parameters; constructing a historical state matrix based on the preprocessed data and the input charging parameters; and extracting the long-term time series dependencies of the historical state matrix through a multi-head attention mechanism, thereby generating a high-dimensional feature vector.
[0006] Preferably, the charging parameters include the capacitor voltage V C (t), the bus voltage U dc (t), and the temperature T(t); the preprocessing process is as follows: calculating the voltage error e(t) between the capacitor voltage V C (t) and the set target voltage V ref and the corresponding error change rate Δe(t).
[0007] Preferably, the historical state matrix Xt-N:t and the eigenvector h t has the following expression: ; h t = TransformerEncoder(X t-N:t ); where N represents the time window length and N = 5.
[0008] Preferably, the long-term return evaluation of the actions generated by the Actor network includes the following process: receiving the eigenvector h encoded by Transformer t entering a fully connected neural network to output the action a t ; using the Q-value function to input the eigenvector h t and the action a t to evaluate the long-term return Q(h t , a t ) of the actions generated by the Actor.
[0009] Preferably, the action a generated by the Actor network t includes the duty cycle adjustment amount ΔD(t) and the current limiting threshold I limit , and the acquisition of the action a t includes the following process: inputting the eigenvector h t into a fully connected neural network, and outputting the action a through the hidden layer activation function ReLU and the output layer activation function Tanh t ; the expression of the action a t is as follows: a t = [ΔD(t), I limit = Tanh(W 3 × ReLU(W 2 × ReLU(W 1 × h t + b 1 ) + b 2 ) + b 3 ); where b 1 to b 3 and W 1 to W 3 are all network parameters.
[0010] Preferably, the optimization of the long-term return includes the following process: optimizing the policy by maximizing the cumulative reward, designing the reward function in combination with the real-time charging effect; based on the obtained reward function, evaluating the value of the actions generated by the Actor network through the Critic network; updating the parameters corresponding to the Critic network and the Actor network by gradient descent to optimize the long-term return.
[0011] Preferably, the reward function includes a charging efficiency reward R 1 , a stability penalty R 2 and an overcharge / undercharge penalty R 3 , and the total reward R = R 1 + R 2 + R 3 ; The charging efficiency reward R 1 , the stability penalty R 2 and the overcharge / undercharge penalty R 3 are expressed as follows: ; ; ; where α, β, and γ respectively represent the corresponding weight coefficients, V ref represents the target voltage, and ε represents the allowable deviation threshold.
[0012] Preferably, the pre-charge control model is trained through meta-learning. The specific training process is as follows: Define the data fluctuation ranges of the target voltage, temperature, and bus voltage in multiple scenarios, and then generate the time-series data for each scenario; Perform a small number of gradient updates and learn the general initialization parameters in each scenario to obtain a set of general parameters; Judge the error range of the general parameters. If the error is too large, fine-tune and update the general parameters and deploy the updated parameters to the real-time control system; Otherwise, maintain the current policy without update.
[0013] Preferably, the upper limit of the difference error range between the temperature in the general parameters and the average value is 10%, and the upper limit of the fluctuation error range of the bus voltage in the general parameters is 5%.
[0014] Compared with the prior art, the beneficial effects of this application are as follows: The DRL-Transformer hybrid framework combines the time-series modeling and dynamic optimization capabilities, effectively improving the voltage control accuracy to ensure that the system can still operate stably under complex working conditions such as sudden temperature changes and bus voltage fluctuations. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic structural diagram of the resonant soft-switching push-pull topology circuit of this application.
[0016] Figure 2 is a schematic diagram of the working process of this application.
[0017] Figure 3 is a schematic diagram of the basic architecture of the pre-charge control model in this application.
[0018] Figure 4It is a logic schematic diagram for pre-charge control of the pre-charge control model in this application.
[0019] Figure 5 It is a process schematic diagram for meta-learning training of the pre-charge control model in this application. Detailed implementation manners
[0020] Next, in combination with the detailed implementation manners, the present application will be further described. It should be noted that in the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0021] In the description of the present application, it should be noted that for orientation terms, if there are terms such as "center", "horizontal", "longitudinal", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc., indicating the orientation and position relationship are based on the orientation or position relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and should not be understood as limiting the specific protection scope of the present application.
[0022] It should be noted that the terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0023] In the present application, unless otherwise clearly defined and limited, the terms "installed", "connected", "connected", "fixed", etc. should be understood in a broad sense. For example, it can be a connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0024] In this application, unless otherwise clearly defined or limited, the first feature being "on" or "under" the second feature may include direct contact between the first and second features, or may include the first and second features not being in direct contact but in contact through additional features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes the first feature being directly above and obliquely above the second feature, or merely indicating that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature includes the first feature being directly below and obliquely below the second feature, or merely indicating that the horizontal height of the first feature is less than that of the second feature.
[0025] The terms "comprising" and "having" and any variations thereof in the description and claims of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0026] For the convenience of understanding the technical solution of this application, the specific structure and working process of the pre-charging hardware architecture of the high-voltage high-power floating capacitor in this application will be described in detail below.
[0027] As Figure 1 shown, it is a schematic diagram of the specific structure of the resonant soft-switching push-pull topology circuit required for pre-charging the high-voltage high-power floating capacitor. This circuit includes two switching tubes Q 1 and Q 2 , a transformer T with a center tap, a full-bridge rectifier circuit (D 1 , D 2 , D 3 , D 4 ) and a filter capacitor C. The primary side of the push-pull transformer T is alternately turned on by the switching tubes Q 1 and Q 2 to form a push-pull structure, generating an alternating current in positive and negative directions. The secondary side rectifies the alternating current into high-voltage direct current through the bridge rectifier circuit, and finally pre-charges the floating capacitor through filtering by the filter capacitor.
[0028] Specifically, as Figure 1 shown, the primary coil of the push-pull transformer is divided into two symmetric windings N 11 and N 12 by the center tap, so as to balance the current and magnetic flux in the primary. The drains of the two switching tubes are respectively connected to one end of N 11 and N 12 , and the power supply terminal is provided by a 12V DC power supply. The control signals are output through the PWM controller to control the switching tubes Q 1 and Q2 alternately conducts. When the switching transistor Q 1 conducts, the current flows from the 12V power supply through the upper half winding of the primary winding of the transformer, passes through the switching transistor Q 1 and is grounded. At this time, the primary winding of the transformer generates magnetic flux, which is transferred to the secondary winding. When the switching transistor Q 2 conducts, the current flows in the opposite direction through the lower half winding of the primary winding, forming an opposite magnetic flux; this bidirectional magnetic flux change helps to improve the utilization rate of the transformer. On the secondary side, the output voltage U 2 is expressed as: U 2 =N 2 / N 1 ×Ui.
[0029] In the formula, N 1 =N 11 =N 12 ; the number of turns of the two windings is the same to ensure that the transformer can provide the same magnetic flux change amount in each switching cycle. U 2 is the voltage output by the secondary of the transformer, and Ui is the voltage input to the push-pull circuit. Then, the high-frequency alternating voltage is converted into direct current through the rectification and filtering circuits. Due to the large turns ratio, the voltage output by the secondary is much higher than the voltage input to the primary.
[0030] The rectification part uses a full-bridge rectifier, which consists of four fast-recovery diodes with high-voltage withstand. When the alternating voltage U 2 of the secondary of the transformer passes through the rectifier bridge, the diodes D 1 and D 4 conduct in the positive half cycle of the voltage, and the diodes D 2 and D 3 conduct in the negative half cycle, ensuring that there is a positive current passing through the output terminal of the rectification circuit in each cycle, thereby converting the alternating current into a unidirectional pulsating direct current. There is still a large ripple voltage in the pulsating direct current output by the rectification circuit. The filter capacitor acts as an energy storage element in the circuit. When the voltage fluctuates, the capacitor can absorb and release charges to smooth the output voltage, reduce the influence of the ripple, and ensure that the pre-charging process of the floating capacitor can proceed smoothly.
[0031] In this circuit, the secondary voltage of the push-pull transformer is connected in parallel with the floating capacitor, and a high-voltage DC power supply is formed through the full-bridge rectifier and the filter capacitor. As the soft-start process progresses, the voltage across the floating capacitor gradually rises to the target value to complete the pre-charging. After the pre-charging is completed, the system enters the normal working state, and the floating capacitor maintains voltage balance, providing stable voltage support for the subsequent five-level topology. In order to ensure the stable start of the resonant soft-switching push-pull topology circuit, this embodiment provides a control method for pre-charging a high-voltage and high-power floating capacitor.
[0032] One preferred embodiment of the present application is as follows Figure 2 and Figure 3 shown, a high-voltage high-power floating capacitor pre-charge control method, applied to the above-mentioned resonant soft-switching push-pull topology circuit, includes the following control steps: constructing a pre-charge control model based on the fusion framework of deep learning and attention mechanism (DRL-Transformer). Taking the real-time collected floating capacitor charging parameters as the input of the pre-charge control model to generate feature vectors. Based on the obtained feature vectors, perform long-term reward evaluation on the actions generated by the Actor network for adjusting the pre-charge process. Optimize the long-term reward according to the real-time charging effect of the floating capacitor and output the optimized actions.
[0033] It can be understood that deep reinforcement learning (DRL) refers to the combination of reinforcement learning and deep learning, which optimizes the reinforcement learning algorithm through neural network deep learning technology. In this embodiment, by integrating deep reinforcement learning with the attention mechanism, combined with the resonant soft-switching push-pull topology circuit, efficient and high-precision capacitor pre-charge can be achieved. The Transformer model architecture uses the Self-Attention structure to replace the RNN network structure commonly used in NLP tasks. Compared with the RNN network structure, its biggest advantage is that it can perform parallel computing. The Transformer model is a neural network that learns context and thus meaning by tracking relationships in sequential data. In this embodiment, the Transformer encoder is used to extract long-time sequence dependencies, generate high-dimensional feature vectors, and use the DDPG algorithm to generate continuous control actions (such as duty cycle adjustment amount, current threshold), and evaluate the long-term reward of the actions through the Critic network to achieve the best pre-charge control of the floating capacitor. Specifically, in this embodiment, the DRL-Transformer hybrid framework combines time series modeling and dynamic optimization capabilities to effectively improve the voltage control accuracy to ensure that the system can still operate stably under complex working conditions such as temperature mutation and bus voltage fluctuation.
[0034] Specifically, as Figure 3 and Figure 4As shown in the figure, the basic structure of the pre-charge control model based on the DRL-Transformer hybrid framework mainly includes an input layer, a Transformer encoder, a DRL network, and an execution layer. The input layer is mainly used for collecting data input and calculating related preprocessing processes. The Transformer encoder can receive the data output by the input layer and perform state encoding and feature extraction to obtain corresponding feature vectors. The DRL network can generate corresponding charging actions by receiving the feature vectors output by the Transformer encoder through the Actor network. At the same time, the DRL network can also evaluate and optimize the long-term charging benefits brought by the charging actions generated by the Actor network through the Critic network. The execution layer can output circuit control signals by combining the real-time charging effect and the optimization result of the long-term charging benefit, and then control the duty cycle and charging current of the resonant soft-switching push-pull topology circuit during the pre-charge process, so that the floating capacitor pre-charge always works in a better or optimal state. For the convenience of understanding, the specific working processes of each part of the pre-charge control model will be described in detail below.
[0035] In this embodiment, as Figure 4 shown, the generation of the feature vector includes the following process: The input layer first preprocesses the input charging parameters. Then, the Transformer encoder constructs a historical state matrix based on the preprocessed data combined with the input charging parameters. Finally, the long-term time series dependence of the historical state matrix is extracted through the multi-head attention mechanism, and then a high-dimensional feature vector is generated.
[0036] Specifically, the charging parameters corresponding to the floating capacitor during the charging process mainly include the capacitor voltage V C (t), the bus voltage U dc (t), and the temperature T(t). Among them, the parameter that has the most obvious response to the charging state of the floating capacitor is the capacitor voltage V C (t). Then, the preprocessing of the charging parameters is mainly to calculate the change state of the capacitor voltage V C (t) during the charging process; that is, to calculate the voltage error e(t) between the capacitor voltage V C (t) and the set target voltage V ref and the corresponding error change rate Δe(t).
[0037] It can be understood that e(t)=V ref -V C (t), Δe(t)=e(t)-e(t - 1). The target voltage V ref is related to the specific performance of the floating capacitor; for the specific value of the target voltage V ref , those skilled in the art can select it according to actual needs.
[0038] It can also be understood that the historical state matrix X t-N:t and the eigenvector h t are expressed as follows: .
[0039] h t = TransformerEncoder(X t-N:t ).
[0040] Where N represents the time window length, and N = 5.
[0041] In this embodiment, as Figure 4 shown, the DRL network adopts an Actor-Critic architecture based on the DDPG algorithm; after receiving the eigenvector h t output by the Transformer encoder, the DRL network can first generate the corresponding action a t through the Actor network. The action a t includes the duty cycle adjustment amount ΔD(t) and the current limiting threshold I limit . Then the DRL network can evaluate the long-term charging benefit of the action a t generated by the Actor network through the Critic network, and optimize according to the evaluation result, and then can output the optimized action a t to the execution layer. The execution layer sends the optimized duty cycle adjustment amount ΔD(t) to the PWM controller to adjust the duty cycle of the switching tubes of the resonant soft-switching push-pull topology circuit; at the same time, the execution layer can also send the optimized current limiting threshold I limit to the synchronous rectification module of the resonant soft-switching push-pull topology circuit, and then limit the charging current I chg (t) of the floating capacitor during the pre-charging process to I chg (t) ≤ I limit .
[0042] It can be understood that the Actor network is a policy gradient algorithm based on policy, which can select behaviors according to probability. The Actor network directly interacts with the environment according to the current policy, and then directly optimizes the current policy according to the rewards obtained after the interaction. The Critic network is a Q-Learning algorithm based on the value function, which is used to judge the behavior score of the Actor network; the update of the Critic network adopts the method of gradient descent. The Critic network directly obtains the interaction between the policy and the environment through the value function, and the rewards obtained from the interaction are used to optimize the current value function, thereby helping the Actor network to update the policy.
[0043] It should be noted that the specific working processes of the Actor network and the Critic network are well known to those skilled in the art, so they will not be elaborated in detail here.
[0044] Specifically, as Figure 4 shown, the long-term reward evaluation of the Actor network for generating actions includes the following process: receiving the feature vector h encoded by the Transformer t and entering a fully connected neural network to output the action a t . Using the Q-value function to input the feature vector h t and the action a t to evaluate the long-term reward Q(h t , a t ) generated by the Actor. The acquisition of the action a t includes the following process: inputting the feature vector h t into a fully connected neural network, and outputting the action a through the hidden layer activation function ReLU and the output layer activation function Tanh t ; the expression of the action a t is as follows: a t =[ΔD(t), I limit =Tanh(W 3 ×ReLU(W 2 ×ReLU(W 1 ×h t +b 1 )+b 2 )+b 3 ); where b 1 to b 3 and W 1 to W 3 are all network parameters.
[0045] In this embodiment, the optimization of the long-term reward by the DRL network includes the following process: optimizing the policy by maximizing the cumulative reward, and designing the reward function in combination with the real-time charging effect. Based on the obtained reward function, the Critic network evaluates the value of the actions generated by the Actor network. The parameters corresponding to the Critic network and the Actor network are updated by gradient descent to optimize the long-term reward.
[0046] It can be understood that the pre-charging process of the floating capacitor mainly considers the charging efficiency, the stability during the charging process, and whether there is overcharging or undercharging behavior. Then, when optimizing the long-term reward, the reward function needs to comprehensively consider the above three situations.
[0047] Specifically, the reward function includes the charging efficiency reward R 1 and the stability penalty R 2And overcharge / undercharge penalty R 3 , the total reward R = R 1 + R 2 + R 3 ; charging efficiency reward R 1 , stability penalty R 2 And overcharge / undercharge penalty R 3 The expressions are as follows: .
[0048] .
[0049] .
[0050] Among them, α, β, and γ respectively represent the corresponding weight coefficients, and ε represents the allowable deviation threshold.
[0051] In this embodiment, as Figure 3 shown, to improve the generalization ability of the pre-charging control model in new scenarios (such as temperature mutation, load fluctuation), the core idea of introducing Meta-Learning is to pre-train the model in multiple simulation scenarios to learn a general strategy, enabling it to quickly adjust the strategy parameters through a small amount of real-time data in new scenarios, so as to adapt to different working conditions. As Figure 5 shown, the specific training process is as follows: Define the data fluctuation ranges of the target voltage V ref , temperature T, and bus voltage U dc in multiple scenarios, and then generate the time-series data in each scenario. Perform a small amount of gradient updates and learn the general initialization parameters in each scenario to obtain a set of general parameters. Judge the error range of the general parameters. If the error is too large, fine-tune and update the general parameters and deploy the updated parameters to the real-time control system.
[0052] It can be understood that the number of samples generated for each scenario can be determined according to actual needs. For example, 1000 - 5000 samples can be generated for each scenario, so as to ensure the diversity of data. The limitation of the error range can also be set according to the actual needs of those skilled in the art. For example, the upper limit of the difference error range between the temperature T(t) and the average value T avg in the general parameters is 10%, and the upper limit of the fluctuation error range of the bus voltage U dc in the general parameters is 5%. The change amount of the fluctuation amplitude of the bus voltage U dc can be represented by ΔU dc .
[0053] Specifically, if |T(t) - T avg | > 10%, and ΔU dcWhen it is > 5%, it indicates that the error range of the general parameters is too large. At this time, a small number of samples can be collected, such as 10 real-time sample parameters D new Perform rapid fine-tuning and update, and then use the updated parameter D new Deploy it to the real-time control system to adjust and control the pre-charging process of the floating capacitor; otherwise, maintain the current strategy without parameter update.
[0054] The above describes the basic principle, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present application. Without departing from the spirit and scope of the present application, the present application will have various changes and improvements, and these changes and improvements all fall within the scope of the present application claimed. The scope of protection required by the present application is defined by the appended claims and their equivalents.
Claims
1. A high-voltage and high-power floating capacitor pre-charging control method, applied to a resonant soft-switching push-pull topology circuit, characterized in that: The control steps include: Construct a pre-charging control model based on a deep learning and attention mechanism fusion framework; The real-time collected charging parameters of the suspended capacitor are used as the input of the pre-charging control model to generate a feature vector; Based on the obtained feature vectors, the long-term benefits of the actions generated by the Actor network for adjusting the precharging process are evaluated; The long-term benefits are optimized based on the real-time charging effect of the suspended capacitor and the optimized actions are output.
2. The high-voltage and high-power floating capacitor pre-charging control method according to claim 1, characterized in that: The generation of feature vectors includes the following process: Preprocessing the input charging parameters; Construct a historical state matrix based on the preprocessed data combined with the input charging parameters; The long-term temporal dependencies of the historical state matrix are extracted through the multi-head attention mechanism to generate a high-dimensional feature vector.
3. The high-voltage and high-power floating capacitor pre-charging control method according to claim 2, characterized in that: The charging parameters include capacitor voltage, bus voltage and temperature; the preprocessing process is as follows: The voltage error between the capacitor voltage and the set target voltage and the corresponding error change rate are calculated.
4. The high-voltage and high-power floating capacitor pre-charging control method according to claim 3, characterized in that: Historical state matrix X t-N:t And the eigenvector h t The expression is as follows: ; h t =TransformerEncoder(X t-N:t ); Among them, V C (t) represents the capacitor voltage, U dc (t) represents bus voltage, T(t) represents temperature, e(t) represents voltage error, Δe(t) represents error change rate, N represents time window length, N=5.
5. The high-voltage and high-power floating capacitor pre-charging control method according to claim 1, characterized in that: Actions generated by the Actor network t The long-term benefit evaluation process is: Receive the Transformer-encoded feature vector h t Enter the fully connected neural network to output action a t ; Use the Q value function to input the feature vector h t and action a t Evaluate the long-term benefit Q(h) of the Actor's generated action t , a t ).
6. The high-voltage and high-power floating capacitor pre-charging control method according to claim 5, characterized in that: Actions generated by the Actor network t Including duty cycle adjustment ΔD(t) and current limit threshold I limit , action a t The acquisition of includes the following process: The feature vector h t Input into the fully connected neural network, output action a through the hidden layer activation function ReLU and the output layer activation function Tanh t ; Action a t The expression is as follows: a t =[ΔD(t),I limit ]=Tanh(W3×ReLU(W2×ReLU(W1×h t +b1)+b2)+b3); Among them, b1 to b3 and W1 to W3 are network parameters.
7. The high-voltage and high-power floating capacitor pre-charging control method according to claim 1, characterized in that: Optimization of long-term returns includes the following processes: By maximizing the cumulative reward optimization strategy, the reward function is designed in combination with the real-time charging effect; Based on the obtained reward function, the value of the action generated by the Actor network is evaluated through the Critic network; The corresponding parameters of the Critic network and the Actor network are updated through gradient descent to optimize long-term benefits.
8. The high-voltage and high-power floating capacitor pre-charging control method according to claim 7, characterized in that: The reward function includes charging efficiency reward R1, stability penalty R2, and overcharge / undercharge penalty R3. The total reward R = R1+R2+R3; The expressions for charging efficiency reward R1, stability penalty R2, and overcharge / undercharge penalty R3 are as follows: ; ; ; Among them, α, β and γ represent the corresponding weight coefficients, V ref represents the target voltage, ε represents the allowable deviation threshold, e(t) represents the voltage error, Δe(t) represents the error change rate, V C (t) represents the capacitor voltage.
9. The high-voltage and high-power suspension capacitor pre-charging control method according to any one of claims 1 to 8, characterized in that: The pre-charging control model is trained through meta-learning. The specific training process is as follows: Define the data fluctuation range of target voltage, temperature and bus voltage in multiple scenarios, and then generate time series data in each scenario; Perform a small number of gradient updates in each scenario and learn a common initialization parameter to obtain a set of common parameters; Determine the error range of the general parameters. If the error is too large, fine-tune and update the general parameters and deploy the updated parameters to the real-time control system; Otherwise, maintain the current policy without updating.
10. The high-voltage and high-power floating capacitor pre-charging control method according to claim 9, characterized in that: The upper limit of the error range of the difference between the temperature in the general parameters and the average value is 10%, and the upper limit of the error range of the fluctuation of the bus voltage in the general parameters is 5%.
Citation Information
Patent Citations
Fuzzy control method and device for high-voltage capacitor charging
CN103457475A
Lithium ion battery health state evaluation method and apparatus, and electronic device
CN118938056A
Unmanned aerial vehicle intelligent charging control method and system based on multi-scene adaptation
CN119527602A
Equalization charging method, device and equipment for power battery pack
CN119765584A
Charging and power supply optimization method and apparatus for charging management system
US20240294086A1