Multi-objective model predictive control method based on tolerance sequential optimization
Patent Information
- Application Number
- CN202611242251.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-17
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]本发明的目的是提供一种精度高、适应性强的基于容差顺序优化的多目标模型预测控制方法,其能够解决传统多目标预测控制中权重系数复杂的问题,同时实现多维变量全局优化及快速动态协调
采用上述结构后,本发明具有如下有益效果:采用将成本函数拆解成单一目标的误差项(即转矩目标项和磁链目标项),以消除了传统方法中成本函数繁琐的权重系数设计,避免了因固定权重系数无法适应多变工况而造成的控制性能下降的问题;同时,根据目标误差项的熵值大小动态调整目标优先级,以便自适应调整转矩与磁链控制优先级分配,实现不同控制周期的目标针对性控制;其中,引入容差阈值对控制误差进行约束,配合自适应顺序寻优机制选择最优矢量,以提高电机控制系统的动态自适应性和全速域下的动态响应与稳态精度,并实现多维变量全局优化及快速动态协调。
Smart Images

Figure CN122824045A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of motor control, and more specifically to a multi-objective model predictive control method based on tolerance sequence optimization. Background Technology
[0002] New energy vehicles (NEVs), as an alternative to traditional gasoline-powered vehicles, have become a mainstream trend in the automotive industry. One of the core technologies of NEVs lies in the electric drive system, which mainly consists of an electric motor and an inverter, and directly determines the vehicle's power, economy, and driving smoothness.
[0003] Currently, the electric drive systems of new energy vehicles mainly adopt traditional PI control methods and model predictive control methods. In the PI control method, the PI control method relies on an accurate mathematical model of the motor, which makes the control accuracy easily affected by parameter changes under dynamic conditions such as high-speed cruising or frequent start-stop. The dynamic response bandwidth is limited, and it is difficult to effectively handle the nonlinear constraints of the inverter.
[0004] In model predictive control, the design of weight coefficients is highly dependent on the parameter tuning process, which is cumbersome and subjective. Furthermore, it cannot achieve optimal global performance under complex and ever-changing driving conditions.
[0005] In addition, the control requirements of electric drive systems for new energy vehicles are to ensure rapid dynamic response while also taking into account optimization of multiple objectives such as reducing torque and flux ripple and suppressing current harmonics. However, the traditional control methods mentioned above are difficult to effectively achieve global optimization and rapid dynamic coordination of multi-dimensional variables.
[0006] In view of this, this application has conducted in-depth research on this basis, resulting in this case. Summary of the Invention
[0007] The purpose of this invention is to provide a high-precision and highly adaptable multi-objective model predictive control method based on tolerance-order optimization, which can solve the problem of complex weight coefficients in traditional multi-objective predictive control, and at the same time achieve global optimization of multi-dimensional variables and rapid dynamic coordination.
[0008] To achieve the above objectives, the solution of this invention is: a multi-objective model predictive control method based on tolerance sequence optimization, applicable to permanent magnet synchronous motors, wherein the permanent magnet synchronous motor is equipped with an electric drive system, the electric drive system including an inverter and a speed loop PI controller, the inverter having eight pre-selected vectors; comprising the following steps: Step 1: Establish the mathematical model of the permanent magnet synchronous motor; Step 2: Based on the mathematical model, obtain the stator flux prediction calculation formula and the electromagnetic torque prediction calculation formula, and establish the cost function of model predictive control; Step 3: Select the optimal vector based on the tolerance-order optimization mechanism; decompose the cost function to obtain the error terms of the torque target and the flux linkage target, then traverse eight vectors to obtain each error term of the torque target and the flux linkage target respectively, and dynamically determine the priority of the torque target and the flux linkage target according to the entropy value of each error term in the torque target and the flux linkage target, which are respectively corresponding to high priority and low priority; then select the pre-selected vector in the high priority that meets the target tolerance threshold as the shortlisted vector and enter it into the high priority set. Select according to the number of shortlisted vectors in the high priority set. If the number is no more than 1, output the corresponding vector as the global optimal vector. Otherwise, select the shortlisted vector in the high priority set that meets the other target tolerance threshold as the candidate vector and enter it into the low priority set. Select according to the number of candidate vectors in the low priority set. If the number is no more than 1, output the corresponding vector as the global optimal vector. Otherwise, select the vector with the smallest sum of the error term values in the high priority set and the low priority set as the global optimal vector. Step 4: The optimal vector output is sent to the inverter to control the on / off state of each switch in the inverter, and then the inverter outputs three-phase current to the permanent magnet synchronous motor.
[0009] In step 1, the mathematical model of the permanent magnet synchronous motor is: , , , ; In the formula, , They represent dq Stator voltage in axial coordinate system , They represent dq Stator current in axial coordinate system Let be the stator resistance of the permanent magnet synchronous motor. , These represent the stator flux linkages. dq Axial components, Indicates permanent magnet flux linkage. , These represent the stator inductance. dq Axial components, This represents the rotor mechanical angular velocity of the permanent magnet synchronous motor. θThis indicates the rotor mechanical angle of the permanent magnet synchronous motor. Indicates electromagnetic torque. This indicates the number of pole pairs of the permanent magnet synchronous motor.
[0010] The inverter has three-phase bridge arms, and the three-phase bridge arms correspond to respectively as follows: a Mutually, b Harmony c The inverter has two phases, and each phase is equipped with two switching transistors. The two switching transistors in each phase's bridge arm are respectively located in the upper bridge arm and the lower bridge arm; wherein, the output voltage calculation formula of the inverter is, In the formula, Indicates the output voltage. Indicates DC voltage. , and These respectively represent the inverter in Mutually, Harmony The switching state of the upper bridge arm switch transistor. Represents the imaginary unit. Represents the base of the natural logarithm in complex form; =0、 =0 and =0 respectively represent the inverter in Mutually, Harmony The upper bridge arm switch of the phase is turned off. =1、 =1 and =1 respectively represent the inverter in Mutually, Harmony The upper bridge arm switch of the phase is turned on; Specifically, a three-bit binary number is obtained based on the switching states of the upper bridge arm switches of the three phases in the inverter. Using control input As a pre-selected vector This indicates the current state of the switch.
[0011] The three-phase voltage is obtained based on the relationship between the switch state and the three-phase voltage of the permanent magnet synchronous motor, and then... Clarke Transformation and Park Transform the three-phase voltage from abc Transformation to three-phase natural coordinate system dq In a two-phase synchronous rotating coordinate system, the relationship is: , , ; in, Clarke Transform into , Park Transform into ; In the formula, , , They represent Phase voltage, Phase voltage, Phase voltage, , They represent Stator voltage in axial coordinate system , They represent dq Stator voltage in axial coordinate system θ This refers to the rotor mechanical angle of the permanent magnet synchronous motor.
[0012] In step 2, the stator voltage equations in the mathematical model of the permanent magnet synchronous motor are discretized using forward Euler discretization to obtain the values of the permanent magnet synchronous motor in the model. dq The formula for predicting current in the axial coordinate system is: , ; In the formula, For the first +1 cycle Predicted current value of the shaft. For the first +1 cycle Predicted current value of the shaft. Indicates the first One cycle Shaft current sampling value, Indicates the first One cycle Shaft current sampling value, The time interval of the sampling period; Based on the current prediction formula and the mathematical model of the permanent magnet synchronous motor, the stator flux prediction calculation formula and the electromagnetic torque prediction calculation formula are obtained, respectively: , , , ; In the formula, Indicates in Predicted electromagnetic torque at time +1 Indicates in Stator flux prediction at time +1 Indicates in +1 time Shaft stator flux linkage prediction value Indicates in +1 time Predicted values of stator flux linkage; The cost function is: In the formula, and They represent in Predicted values of electromagnetic torque and stator flux linkage at time +1 This represents the torque reference value output by the outer loop PI controller for the rotational speed. Indicates the stator flux linkage reference value. Indicates stator flux linkage. represents the weighting coefficient of the magnetic flux term.
[0013] Step 3 includes the following steps; Step 3-1, Decoupling the cost function; the cost function is: In the formula, and They represent in Predicted values of electromagnetic torque and stator flux linkage at time +1 This represents the torque reference value output by the outer loop PI controller for the rotational speed. Indicates the stator flux linkage reference value. Indicates stator flux linkage. By structurally decomposing the cost function using the weighting coefficients of the flux linkage term, the error terms for the torque target and flux linkage target are obtained as follows: , In the formula, The error term representing the torque target, The error term representing the magnetic flux target, This represents the torque output value of the speed loop PI controller. Indicates the stator flux linkage reference value; Step 3-2: Determine priorities; First, normalize the error terms of the torque target and the flux linkage target. The normalization formula is as follows: , In the formula, This represents the normalized error of the torque. This represents the normalized error of the magnetic flux linkage. express These correspond to seven voltage vectors respectively; Then, the entropy values of the error terms for the torque target and the flux linkage target are calculated using the information entropy formula, which is: , In the formula, and The entropy values of the error terms for the torque target and the flux linkage target are respectively represented; then, the target with the larger entropy value is selected as the high priority, and the other target is selected as the low priority. Step 3-3, Target Sorting: Assign error item ranking values to the torque target from smallest to largest, and assign error item ranking values to the flux linkage target from smallest to largest. Steps 3-4: Determine the globally optimal vector; The cost function is decomposed to obtain the error terms of the torque target and the flux linkage target. Then, the eight vectors are traversed to obtain each error term of the torque target and the flux linkage target. Then, the priority of the torque target and the flux linkage target is dynamically determined according to the entropy value of each error term in the torque target and the flux linkage target, and they are respectively assigned as high priority and low priority. Then, the pre-selected vectors that meet the target tolerance threshold in the high-priority set are selected as the shortlisted vectors and added to the high-priority set. If the number of shortlisted vectors in the high-priority set is no more than one, and there is only one shortlisted vector, then the shortlisted vector is taken as the globally optimal vector. If there are zero shortlisted vectors, then the vector that minimizes the error term in the high-priority set among the eight pre-selected vectors is selected as the globally optimal vector. If the number of shortlisted vectors in the high-priority set is greater than one, then the vectors that meet another target tolerance threshold in the high-priority set are selected as the globally optimal vector. The shortlisted vectors are used as candidate vectors and enter the set of secondary priority vectors. If the number of candidate vectors in the set of secondary priority vectors is no more than 1, and there is 1 candidate vector, then the candidate vector is used as the global optimal vector. If there are 0 candidate vectors, then the vector that minimizes the error term in the secondary priority vector among the shortlisted vectors is selected as the global optimal vector. If the number of candidate vectors in the set of secondary priority vectors is greater than 1, then the vector with the smallest sum of the error term ranking values in the high priority set and the secondary priority set is selected as the global optimal vector.
[0014] In steps 3-4, the two target tolerance thresholds correspond to the torque tolerance threshold and the flux tolerance threshold. The two target tolerance thresholds are adaptively adjusted using the following algorithm, which specifically includes the following steps; Step S1: Establish the Actor network policy function; The correspondence between the operating state of the electric drive system and the output tolerance threshold is represented by an Actor network strategy function, which is: In the formula, This represents a function that retrieves the action. Represents an Actor network. Indicates the parameters of the Actor network; This represents the input motor environment state vector, where, In the formula, These represent the actual torque error and flux linkage error, respectively. Indicates the actual rotational speed; Step S2: Establish and update the mean square error loss function of the Critic network; wherein the mean square error loss function of the Critic network is: In the formula, For small batch sample size, This represents the loss function value of the Critic network. Indicates the parameters of the Critic network; Represents the Critic network, where, ; Indicates the first The target action value corresponding to the sample, obtained from the sampling of the first sample. The sample is The objective value function is: ; In the formula, Indicates the discount factor. Represents the target Actor network; Represents the target Critic network. and These represent the parameters of the target Actor network and the target Critic network, respectively. Indicates the first The current state of each sample; Indicates the execution of an action The environment then transitions to the state of the next moment; Indicates the first The action performed in the current state of a sample is usually output by the Actor network at that time; Indicates the execution of an action The immediate reward value returned by the environment is then used to evaluate the merits of the action. Step S3: The Actor network continuously optimizes the output policy using a policy gradient update method; The formula for the policy gradient update is: , ; In the formula, Indicates the parameters of the Actor network. The gradient; The action value function is expressed with respect to the action variable. The gradient; This represents the continuous actions output by the Actor network based on the current state; This represents the objective function for policy performance. Step S4: The Actor network parameters and the Critic network parameters are updated using gradient ascent and gradient descent methods, respectively, with the following formulas: , ; In the formula, and These are the learning rates for the Actor network and the Critic network, respectively.
[0015] Step S5: Update the data from step S4. Substituting this into the Actor network policy function in step S1, we obtain the optimized tolerance threshold.
[0016] In step S4, the target network uses a soft update method, specifically, , ; In the formula, Denotes the target network soft update coefficient, and .
[0017] Before step S1, an experienced replay pool is established, defined as: In the formula, This represents the system state at the next moment after performing the current tolerance threshold tuning action. With the above structure, the present invention has the following beneficial effects: by decomposing the cost function into single-objective error terms (i.e., torque target term and flux target term), the cumbersome weight coefficient design of the cost function in the traditional method is eliminated, avoiding the problem of control performance degradation caused by the inability of fixed weight coefficients to adapt to changing operating conditions; at the same time, the target priority is dynamically adjusted according to the entropy value of the target error term, so as to adaptively adjust the priority allocation of torque and flux control, and realize targeted control of different control cycles; wherein, a tolerance threshold is introduced to constrain the control error, and the optimal vector is selected in conjunction with an adaptive sequential optimization mechanism, so as to improve the dynamic adaptability of the motor control system and the dynamic response and steady-state accuracy in the full speed domain, and realize multi-dimensional variable global optimization and rapid dynamic coordination. Attached Figure Description
[0018] Figure 1 This is a circuit topology diagram of the permanent magnet synchronous motor and inverter in this invention.
[0019] Figure 2 This is a block diagram illustrating the control principle of the multi-objective model predictive control method of the present invention.
[0020] Figure 3 This is a flowchart of the tolerance-based sequential optimization mechanism in this invention.
[0021] Figure 4 In order to be in Figure 2 The flowchart of the DDPG algorithm is added to the basic algorithm. Detailed Implementation
[0022] To further explain the technical solution of the present invention, the present invention will be described in detail below through specific embodiments.
[0023] A multi-objective model predictive control method based on tolerance sequence optimization is applicable to permanent magnet synchronous motors. This method is based on the control of common electric drive systems for permanent magnet synchronous motors. In this embodiment, the permanent magnet synchronous motor can be an embedded permanent magnet synchronous motor.
[0024] like Figures 1-3 As shown, the above-mentioned electric drive system includes an inverter, which is a conventional two-level three-phase inverter. The inverter includes three-phase bridge arms, which correspond to the following... Mutually, Harmony Each phase has two switching transistors, and the two switching transistors in each phase arm are located in the upper and lower arms respectively. For example: upper arm and lower bridge arm The on and off states of the switching transistors are represented by binary numbers "1" and "0" respectively. During operation, in order to prevent short circuits from burning out the inverter, the two switching transistors in each phase arm cannot be turned on at the same time.
[0025] Furthermore, the aforementioned permanent magnet synchronous motor serves as the output load of the inverter. The formula for calculating the pre-selected vector (i.e., the input voltage vector) in this inverter is as follows: In the formula, Indicates the output voltage. Indicates DC voltage. , and These represent the inverters in... Mutually, Harmony The switching state of the upper bridge arm switch transistor. Represents the imaginary unit. Represents the base of the natural logarithm in complex form; =0、 =0 and =0 respectively represent the inverter in Mutually, Harmony The upper bridge arm switch is turned off, and the lower bridge arm switch is turned on. =1、 =1 and =1 respectively represent the inverter in Mutually, Harmony The upper bridge arm switch is turned on, and the lower bridge arm switch is turned off.
[0026] Therefore, based on the switching states of the upper arm switches of the three phases in the inverter, a three-bit binary number is obtained. ,For example , In this embodiment, there are eight pre-selected vectors, numbered... =1~8 indicates that control input is used. As a pre-selected vector This represents the current state of the switch; for example, when... When, it indicates the upper bridge arm Switch cut off, upper bridge arm and All switching transistors are turned on; simultaneously, the switching transistors of each lower bridge arm are in the opposite state, i.e., the lower bridge arm... The switch is turned on, and the lower bridge arm is activated. and The switching transistor is off.
[0027] In addition, the above-mentioned control system also includes an electromagnetic torque and stator flux estimation module, a speed loop PI controller and a model predictive controller. The input quantities of each of these can be obtained by conventional means in the art or by the methods described below, so they will not be described in detail here.
[0028] In this embodiment, the multi-objective model predictive control method includes the following steps.
[0029] Step 1: Establish the mathematical model of the permanent magnet synchronous motor, which includes the stator voltage equation, the stator flux linkage equation, and the electromagnetic torque state-space equation.
[0030] To elaborate, the mathematical model described above is as follows: (1), , (2), (3); In the formula, , They represent dq Stator voltage in axial coordinate system , They represent dq Stator current in axial coordinate system Let be the stator resistance of the permanent magnet synchronous motor. , These represent the stator flux linkages. dq Axial components, Indicates permanent magnet flux linkage. , These represent the stator inductance. dq Axial components, This represents the rotor mechanical angular velocity of the permanent magnet synchronous motor. θ This indicates the rotor mechanical angle of the permanent magnet synchronous motor. Indicates electromagnetic torque. This indicates the number of pole pairs of the permanent magnet synchronous motor.
[0031] Step 2, Model Predictive Torque Control Based on Euler Discretization: The stator flux prediction calculation formula and electromagnetic torque prediction calculation formula are obtained through the mathematical model of the permanent magnet synchronous motor in Step 1, and the cost function of model predictive control is established.
[0032] To elaborate, such as Figure 2 As shown, step 2 includes the following steps.
[0033] Step 2-1, Obtain the current prediction formula: Perform conventional forward Euler discretization on the differential term of formula (1) in the mathematical model of the permanent magnet synchronous motor, i.e. (4), and then simplified to obtain the permanent magnet synchronous motor in dq The formula for current prediction in the axial coordinate system is as follows.
[0034] (5), (6); In the formula, For the first +1 cycle Predicted current value of the shaft. For the first +1 cycle Predicted current value of the shaft. Indicates the first One cycle Shaft current sampling value, Indicates the first One cycle Shaft current sampling value, This is the time interval of the sampling period.
[0035] Furthermore, in the current prediction formula, , The calculation process is as follows: Based on the relationship between the inverter's switching state and the three-phase voltage of the permanent magnet synchronous motor, the three-phase voltage is obtained, and then... Clarke Transformation and Park Transformation from three-phase voltage abc Transformation to three-phase natural coordinate system dq In a two-phase synchronous rotating coordinate system.
[0036] To elaborate, the relationship between the switching state of the inverter and the three-phase voltage is as follows: , , ; In the formula, , , They represent Phase voltage, Phase voltage, Phase voltage.
[0037] in, Clarke Transform into In the formula, , They represent Stator voltage in axial coordinate system.
[0038] Park Transform into In the formula, , They represent dq Stator voltage in axial coordinate system θ This refers to the rotor mechanical angle of the permanent magnet synchronous motor.
[0039] Step 2-2: Obtain the stator flux linkage and electromagnetic torque prediction calculation formulas: Based on the current prediction formula in Step 2-1 and the permanent magnet synchronous motor mathematical model in Step 1, the above model prediction controller can derive the stator flux linkage and electromagnetic torque prediction calculation formulas. That is, by substituting formulas (5) and (6) into formulas (2) and (3), the formulas can be obtained.
[0040] The formulas for predicting and calculating stator flux linkage and electromagnetic torque are as follows: (7), , (8), (9); In the formula, Indicates in Predicted electromagnetic torque at time +1 Indicates in Stator flux prediction at time +1 Indicates in +1 time Shaft stator flux linkage prediction value Indicates in +1 time Predicted values of stator flux linkage.
[0041] Step 2-3: Establish the cost function of model predictive control: To eliminate errors caused by system computation delay and DSP computation delay, after considering one-beat delay compensation, the established cost function of model predictive control is as follows: (10); In the formula, and They represent in Predicted values of electromagnetic torque and stator flux linkage at time +2 This represents the torque reference value output by the outer loop PI controller for the rotational speed. Indicates the stator flux linkage reference value. Indicates stator flux linkage; This is the weighting coefficient for the flux linkage term, used to balance the magnitude difference between torque and flux linkage.
[0042] It is worth mentioning that this embodiment uses... Substituting the numerical value into the above prediction calculation formula, we can obtain... Similarly, we can obtain .
[0043] Step 3: Select the optimal vector based on the tolerance-order optimization mechanism.
[0044] The cost function is decomposed to obtain error terms for torque and flux linkage targets. Then, eight pre-selected vectors are traversed to obtain eight error terms for torque and eight error terms for flux linkage targets. Subsequently, the priorities of torque and flux linkage targets are dynamically determined based on the entropy values of each error term, corresponding to high priority and low priority, respectively. Ranking values are then assigned to torque targets and flux linkage targets from smallest to largest based on their error term values.
[0045] Then, the pre-selected vectors that meet the target tolerance threshold in the high priority set are selected as the shortlisted vectors and enter the high priority set. If the number of shortlisted vectors in the high priority set does not exceed 1, and if there are no shortlisted vectors in the high priority set, the vector that minimizes the error term of the high priority target among the eight pre-selected vectors is selected as the global optimal vector. If there is 1 shortlisted vector in the high priority set, the shortlisted vector is selected as the global optimal vector.
[0046] If the number of shortlisted vectors in the high-priority set is greater than 1, then the shortlisted vectors in the high-priority set that satisfy the other target tolerance threshold are entered into the secondary-priority set as candidate vectors. If the number of candidate vectors in the secondary-priority set is no more than 1, and there are no candidate vectors in the secondary-priority set, then the vector that minimizes the error term of the secondary-priority target is selected from the shortlisted vectors in the high-priority set as the global optimal vector. If there is 1 candidate vector in the secondary-priority set, then that candidate vector is selected as the global optimal vector. If the number of candidate vectors in the secondary-priority set is greater than 1, then the vector with the smallest sum of the error term ranking values in the high-priority set and the error term ranking values in the secondary-priority set is selected as the global optimal vector.
[0047] To elaborate, such as Figure 3 As shown, step 3 includes the following steps.
[0048] Step 3-1: Decoupling the cost function; decompose the cost function in Step 2-3 to obtain the error terms of a single objective, namely the error terms of the torque objective and the error terms of the flux linkage objective, thereby eliminating the weighting coefficients in the cost function.
[0049] To elaborate, the error terms for the torque target term and the flux linkage target term are as follows: , (11); In the formula, The error term representing the torque target, The error term representing the magnetic flux target, This indicates the torque output value of the speed loop PI controller. This indicates the stator flux linkage reference value.
[0050] Step 3-2: Determine the priority order of the targets; traverse the eight pre-selected vectors to obtain the error terms of the eight torque targets and the eight flux targets respectively. Dynamically determine the priority order of the multiple targets based on the entropy value of the target error terms. That is, use the information entropy formula to calculate the entropy value of the error terms of the torque targets and the flux targets respectively, and then select the target with the larger entropy value as the high priority and the other target as the low priority.
[0051] To elaborate, in order to eliminate the influence of dimensions, error normalization is performed on the error terms of the torque target and the flux linkage target respectively. The formula for this normalization is as follows: , In the formula, This represents the normalized error of the torque. This represents the normalized error of the magnetic flux linkage; express These correspond to eight voltage vectors respectively.
[0052] Furthermore, the above formula for information entropy is: , (12), where, and These represent the entropy values of the error terms for the torque target and the flux linkage target, respectively.
[0053] The larger the entropy value, the more discrete and fluctuating the target error distribution, indicating greater uncertainty and a greater impact on system performance under the current operating conditions. Therefore, this embodiment dynamically determines the target priority based on the entropy value. For example, in a certain control cycle, when... > When the torque target fluctuates more drastically, it indicates that the torque target should be given high priority, and the flux linkage target should be given secondary priority; conversely, when... > In this case, the flux linkage target is given high priority, and the torque target is given low priority.
[0054] It is worth mentioning that since information entropy can reflect the uncertainty and volatility of system variables, if the error of a certain target fluctuates greatly, it indicates that the current control state of the target is relatively unstable and has a more significant impact on the overall performance of the system. Therefore, it is given a higher optimization priority in this embodiment.
[0055] Step 3-3: Target sorting.
[0056] Assign ranking values to the error terms of the torque target from smallest to largest, and assign ranking values to the error terms of the flux linkage target from smallest to largest.
[0057] For example, pre-selected vectors If the error term of the magnetic flux target under the action is minimized, then the pre-selected vector... The ranking value is 1, while the preselected vector If the error term of the magnetic flux target under action is suboptimal, then the preselected vector The ranking value is 2, and so on, so that the ranking values of the flux linkage target and torque target corresponding to each pre-selected vector can be obtained.
[0058] Steps 3-4: Determine the globally optimal vector.
[0059] First, among the high-priority targets, the pre-selected vectors that meet the target tolerance threshold are selected as shortlisted vectors and added to the high-priority set. Then, the number of shortlisted vectors in the high-priority set is determined. If the number of shortlisted vectors is no more than 1 (which can be 0 or 1), then the shortlisted vector is taken as the global optimal vector. If the number of shortlisted vectors is 0, then the vector that minimizes the error term of the high-priority targets among the eight pre-selected vectors is selected as the global optimal vector. That is, among the eight pre-selected vectors, the vector with the smallest error ranking value among the high-priority targets is selected.
[0060] If there are more than one shortlisted vector, the shortlisted vectors in the high-priority set that satisfy the tolerance threshold of the other objective are added to the low-priority set as candidate vectors. Then, the number of candidate vectors in the low-priority set is determined. If there is no more than one candidate vector (which can be 0 or 1), this candidate vector is selected as the global optimal vector. If there are zero candidate vectors, the vector that minimizes the error term of the low-priority objective among the shortlisted vectors in the high-priority set is selected as the global optimal vector. If there are more than one candidate vector, the vector with the smallest sum of the error term ranking values in the high-priority set and the low-priority set is selected as the global optimal vector.
[0061] For example, taking torque as a high priority and flux linkage as a second priority, let's illustrate this with a pre-selected vector. , , , The error terms corresponding to the respective torque targets each satisfy the torque tolerance threshold. Therefore, there are four pre-selected vectors that satisfy the torque tolerance threshold. These four pre-selected vectors are then used as inbound vectors and added to the high-priority set. Since there are four inbound vectors in the high-priority set, it is determined whether the flux linkage target error terms corresponding to these four inbound vectors satisfy the flux linkage tolerance threshold. In this embodiment, the inbound vectors... , If the flux linkage tolerance threshold is met, then the included vector , As a candidate vector, it enters the secondary priority set. Since there are two candidate vectors in the secondary priority set... , Then calculate the candidate vector. The candidate vector is calculated similarly by summing the ranking values of the error terms for the torque target and the flux linkage target. The candidate vector is the sum of the ranking values of the error terms for the corresponding torque target and the error terms for the flux linkage target. The calculated value is less than the candidate vector. The calculated value, therefore the candidate vector As the globally optimal vector.
[0062] To illustrate further, let's take torque as a high priority and flux linkage as a second priority example, and illustrate with the pre-selected vector. If the corresponding torque target error term satisfies the torque tolerance threshold, meaning there is one pre-selected vector that satisfies the torque tolerance threshold, then the pre-selected vector will be... As the globally optimal vector.
[0063] To illustrate further, let's take torque as a high priority and flux linkage as a low priority. In the torque target, none of the eight pre-selected vectors meet the torque tolerance threshold. Therefore, the high-priority set contains zero vectors. We then need to determine the torque target error term corresponding to each of the eight pre-selected vectors. In this example, the pre-selected vectors... The corresponding torque target error term is minimized (i.e., the error term ranking value is minimized), and the pre-selected vector is... As the globally optimal vector.
[0064] To illustrate further, let's take torque as a high priority and flux linkage as a second priority example, and illustrate the pre-selection vector. , , , The error terms corresponding to the respective torque targets each satisfy the torque tolerance threshold. Therefore, there are four pre-selected vectors that satisfy the torque tolerance threshold. These four pre-selected vectors are then used as inclusion vectors and added to the high-priority set. Since there are four inclusion vectors in the high-priority set, it is determined whether the flux linkage target error terms corresponding to these four inclusion vectors satisfy the flux linkage tolerance threshold. In this embodiment, none of the four inclusion vectors satisfy the flux linkage tolerance threshold. Therefore, it is determined whether the flux linkage target error terms corresponding to the four inclusion vectors satisfy the flux linkage tolerance threshold. If the corresponding magnetic flux linkage target error term is minimized, then the included vector is selected. As the globally optimal vector.
[0065] Furthermore, torque tolerance thresholds and flux linkage tolerance thresholds are employed to achieve a sequential optimization mechanism in the multi-objective optimization process. Since the tolerance threshold directly affects the selection range of the pre-selected vector, its reasonable setting plays a decisive role in control performance. However, when the operating conditions of the permanent magnet synchronous motor change, a fixed tolerance threshold cannot simultaneously consider dynamic performance, steady-state ripple, and system robustness, easily leading to a decline in control performance. Therefore, in this embodiment, a reinforcement learning environment is constructed with the system operating state as input and control performance indicators as optimization objectives. The DDPG algorithm is selected as the policy learning framework. By establishing a policy network and a value network structure, the current system state variables (including torque error, flux linkage error, current, speed, etc.) are input into the policy network, and the corresponding torque tolerance threshold and flux linkage tolerance threshold are output. Through interaction with the environment, the policy network parameters are continuously updated online, thereby achieving adaptive adjustment of the tolerance threshold.
[0066] To elaborate, such as Figure 4 As shown, the torque tolerance threshold and flux tolerance threshold are adaptively adjusted using the following algorithm, which includes the following steps.
[0067] Step S1: Establish the Actor network policy function.
[0068] The Actor network adaptively adjusts the tolerance thresholds for torque error and flux linkage error based on the current operating state of the electric drive system, achieving online dynamic adjustment of the tolerance thresholds. The relationship between the operating state of the electric drive system and the output tolerance thresholds is represented by the Actor network strategy function, specifically: (13), where, This represents the function that obtains the action, that is, the specific value of the action obtained through the Actor network policy function; Represents an Actor network. This represents the parameters of the Actor network.
[0069] in, This represents the input motor environment state vector. In this embodiment, the current operating state of the electric drive system is used as the input to the reinforcement learning state space, defined as follows: (14), where, These represent the actual torque error and flux linkage error, respectively. This indicates the actual rotational speed. , They represent dq Stator current in axial coordinate system.
[0070] Furthermore, the aforementioned motor environmental state vector includes actual torque error, actual flux linkage error, and qThe stator current and actual rotational speed in the axial coordinate system are five variables in total. Among them, the aforementioned state variables (i.e., It can comprehensively reflect the dynamic characteristics of the system, the operating load, and the current control deviation, providing complete operational information for the reinforcement learning agent.
[0071] Step S2: Establish the Critic network to minimize the mean squared error loss function and update it.
[0072] To elaborate, the Critic network minimizes the mean squared error loss function as follows: (15); In the formula, For small batch sample size, This represents the loss function value of the Critic network. This represents the parameters of the Critic network.
[0073] in, This represents a Critic network. In this embodiment, a Critic network is used to evaluate the value of the current state-action pair. Its state-action value function is defined as: (16).
[0074] Furthermore, regarding the sampled first... Sample The objective value function is: (17), where, Indicates the first The target action value corresponding to each sample is used as the supervision target during the training of the Critic network. Indicates the discount factor. Represents the target Actor network; Represents the target Critic network. and These represent the parameters of the target Actor network and the target Critic network, respectively. Indicates the first The current state of each sample, that is, the environmental state observed by the DDPG agent before performing the action; Indicates the execution of an action Then, the environment transitions to the state of the next moment; Indicates the first The action performed in the current state of a sample is usually output by the Actor network at that time; This indicates that the DDPG agent performs an action. The environment then returns an immediate reward value, which is used to evaluate the merits of the action.
[0075] Step S3: The Actor network continuously optimizes the output policy using a policy gradient update method.
[0076] To elaborate, the formula for the policy gradient update mentioned above is: (18) (19); In the formula, Indicates the parameters of the Actor network. The gradient; The action value function is expressed with respect to the action variable. The gradient; This represents the continuous actions output by the Actor network based on the current state; The objective function for policy performance is the expected cumulative reward that the DDPG agent will obtain under the current policy. This is achieved by backpropagating the gradients of the Critic network with respect to actions to the Actor network parameters through the chain rule, updating them in the direction that increases the expected cumulative reward.
[0077] It should be noted that, in this embodiment, for ease of description, the formula above is modified. abbreviated as .
[0078] In step S4, the Actor network parameters and the Critic network parameters are updated using gradient ascent and gradient descent methods, respectively.
[0079] To elaborate, the specific formula mentioned above is: (20) (twenty one); In the formula, and These are the learning rates for the Actor network and the Critic network, respectively.
[0080] Furthermore, in step S2, formula (16) Substitute into formula (19) for update and optimization, and then... Substituting into formula (21), we can obtain the updated result. Updated in step S4 Substitute it into formula (16), that is, substitute it into... In China, The value is updated to maximize the output of the Critic network. This value is used to assist in updating the main network parameters in order to obtain the optimal tolerance threshold.
[0081] Preferably, to improve training stability, the target network adopts a soft update method: (twenty two), ; In the formula, Denotes the target network soft update coefficient, and .
[0082] In this process, formulas (22) and (23) are substituted into formula (17) to continuously update the target value function, thereby updating... This further assists in updating the main network parameters to obtain the optimal tolerance threshold. Furthermore, the soft update method avoids drastic changes in the target value, improving the stability of the tolerance threshold learning process.
[0083] Thus, the formula obtained above... Substitute into step S1 In this way, the optimized tolerance threshold can be obtained.
[0084] As a preferred approach, this embodiment establishes an experience replay pool to randomly sample small batches of samples from the experience replay pool during training, thereby reducing the temporal correlation between samples and improving the stability of network training.
[0085] To elaborate, the experience replay pool is defined as: In the formula, This represents the system state at the next moment after performing the current tolerance threshold tuning action. The aforementioned experience replay pool is used to store historical sample data for training.
[0086] Furthermore, the tolerance threshold generation strategy using an Actor network is defined as follows: In the formula, and These represent the torque tolerance threshold and the flux linkage tolerance threshold, respectively.
[0087] It should be noted that the DDPG agent belongs to the continuous action space reinforcement learning algorithm based on the Actor-Critic structure. Its core idea is to use the Actor network (policy network) to output continuous control actions and use the Critic network (value network) to evaluate the value of the current action.
[0088] Furthermore, this embodiment includes a reward function, specifically, In the formula, This represents the reward value of the reward function. and These represent the system ratings for torque and flux linkage, respectively.
[0089] Specifically, when the error terms of the torque target and flux target decrease, the reward value increases, thereby driving the DDPG agent to continuously optimize the tolerance output strategy.
[0090] It should be noted that, in order to guide the reinforcement learning agent to learn the optimal tolerance strategy, the above-mentioned reward function is constructed in this embodiment. This reward function takes into account both torque tracking performance and flux linkage control performance.
[0091] In this embodiment, after extensive offline training, the Actor network can learn the mapping relationship of the optimal tolerance threshold under different system states based on the obtained optimal network parameters. The optimal tolerance threshold is then input into the subsequent sequential optimization mechanism (i.e., input into steps 3-4), which can automatically adjust the candidate vector screening range according to the current operating conditions, thereby achieving global optimization under different loads, different speeds, and different disturbance conditions.
[0092] This invention discloses a multi-objective model predictive control method based on tolerance sequence optimization. Taking the real-time system states such as torque error, flux error, current, and speed as inputs, the method outputs and updates the torque tolerance threshold and flux tolerance threshold online through the DDPG algorithm. Based on the target tolerance threshold and the ranking of the error terms corresponding to the vector, combined with an adaptive sequential optimization mechanism, the vector is filtered layer by layer to obtain the globally optimal vector, thereby improving the stability and control performance of the electric drive system.
[0093] The above description is only a preferred embodiment of this invention. Any equivalent changes and modifications made within the scope of the claims of this invention shall fall within the scope of the claims of this invention.
Claims
1. A multi-objective model predictive control method based on tolerance-sequential optimization, applicable to permanent magnet synchronous motors, wherein the permanent magnet synchronous motor is equipped with an electric drive system, the electric drive system including an inverter and a speed loop PI controller, the inverter having eight pre-selected vectors; characterized in that, Includes the following steps: Step 1: Establish the mathematical model of the permanent magnet synchronous motor; Step 2: Based on the mathematical model, obtain the stator flux prediction calculation formula and the electromagnetic torque prediction calculation formula, and establish the cost function of model predictive control; Step 3: Select the optimal vector based on the tolerance-order optimization mechanism; decompose the cost function to obtain the error terms of the torque target and the flux linkage target, then traverse eight vectors to obtain each error term of the torque target and the flux linkage target respectively, and dynamically determine the priority of the torque target and the flux linkage target according to the entropy value of each error term in the torque target and the flux linkage target, which are respectively corresponding to high priority and low priority; then select the pre-selected vector in the high priority that meets the target tolerance threshold as the shortlisted vector and enter it into the high priority set. Select according to the number of shortlisted vectors in the high priority set. If the number is no more than 1, output the corresponding vector as the global optimal vector. Otherwise, select the shortlisted vector in the high priority set that meets the other target tolerance threshold as the candidate vector and enter it into the low priority set. Select according to the number of candidate vectors in the low priority set. If the number is no more than 1, output the corresponding vector as the global optimal vector. Otherwise, select the vector with the smallest sum of the error term values in the high priority set and the low priority set as the global optimal vector. Step 4: The optimal vector output is sent to the inverter to control the on / off state of each switch in the inverter, and then the inverter outputs three-phase current to the permanent magnet synchronous motor.
2. The multi-objective model predictive control method based on tolerance-order optimization according to claim 1, characterized in that: In step 1, the mathematical model of the permanent magnet synchronous motor is: , , , ; In the formula, , They represent dq Stator voltage in axial coordinate system , They represent dq Stator current in axial coordinate system Let be the stator resistance of the permanent magnet synchronous motor. , These represent the stator flux linkages. dq Axial components, Indicates permanent magnet flux linkage. , These represent the stator inductance. dq Axial components, This represents the rotor mechanical angular velocity of the permanent magnet synchronous motor. θ This indicates the rotor mechanical angle of the permanent magnet synchronous motor. Indicates electromagnetic torque. This indicates the number of pole pairs of the permanent magnet synchronous motor.
3. The multi-objective model predictive control method based on tolerance-order optimization according to claim 1, characterized in that: The inverter has three-phase bridge arms, and the three-phase bridge arms correspond to respectively as follows: a Mutually, b Harmony c The inverter has two phases, and each phase is equipped with two switching transistors. The two switching transistors in each phase's bridge arm are respectively located in the upper bridge arm and the lower bridge arm; wherein, the output voltage calculation formula of the inverter is, In the formula, Indicates the output voltage. Indicates DC voltage. , and These respectively represent the inverter in Mutually, Harmony The switching state of the upper bridge arm switch transistor. Represents the imaginary unit. Represents the base of the natural logarithm in complex form; =0、 =0 and =0 respectively represent the inverter in Mutually, Harmony The upper bridge arm switch of the phase is turned off. =1、 =1 and =1 respectively represent the inverter in Mutually, Harmony The upper bridge arm switch of the phase is turned on; Specifically, a three-bit binary number is obtained based on the switching states of the upper bridge arm switches of the three phases in the inverter. Using control input As a pre-selected vector This indicates the current state of the switch.
4. The multi-objective model predictive control method based on tolerance-order optimization according to claim 3, characterized in that: The three-phase voltage is obtained based on the relationship between the switch state and the three-phase voltage of the permanent magnet synchronous motor, and then... Clarke Transformation and Park Transform the three-phase voltage from abc Transformation to three-phase natural coordinate system dq In a two-phase synchronous rotating coordinate system, the relationship is: , , ; in, Clarke Transform into , Park Transform into ; In the formula, , , They represent Phase voltage, Phase voltage, Phase voltage, , They represent Stator voltage in axial coordinate system , They represent dq Stator voltage in axial coordinate system θ This refers to the rotor mechanical angle of the permanent magnet synchronous motor.
5. The multi-objective model predictive control method based on tolerance-order optimization according to claim 4, characterized in that: In step 2, the stator voltage equations in the mathematical model of the permanent magnet synchronous motor are discretized using forward Euler discretization to obtain the values of the permanent magnet synchronous motor in the model. dq The formula for predicting current in the axial coordinate system is: , ; In the formula, For the first +1 cycle Predicted current value of the shaft. For the first +1 cycle Predicted current value of the shaft. Indicates the first One cycle Shaft current sampling value, Indicates the first One cycle Shaft current sampling value, The time interval of the sampling period; Based on the current prediction formula and the mathematical model of the permanent magnet synchronous motor, the stator flux prediction calculation formula and the electromagnetic torque prediction calculation formula are obtained, respectively: , , , ; In the formula, Indicates in Predicted electromagnetic torque at time +1 Indicates in Stator flux prediction at time +1 Indicates in +1 time Shaft stator flux linkage prediction value Indicates in +1 time Predicted values of stator flux linkage; The cost function is: In the formula, and They represent in Predicted values of electromagnetic torque and stator flux linkage at time +1 This represents the torque reference value output by the outer loop PI controller for the rotational speed. Indicates the stator flux linkage reference value. Indicates stator flux linkage. represents the weighting coefficient of the flux linkage term.
6. The multi-objective model predictive control method based on tolerance-order optimization according to claim 5, characterized in that: Step 3 includes the following steps; Step 3-1, Decoupling the cost function; the cost function is: In the formula, and They represent in Predicted values of electromagnetic torque and stator flux linkage at time +1 This represents the torque reference value output by the outer loop PI controller for the rotational speed. Indicates the stator flux linkage reference value. Indicates stator flux linkage. By structurally decomposing the cost function using the weighting coefficients of the flux linkage term, the error terms for the torque target and flux linkage target are obtained as follows: , In the formula, The error term representing the torque target, The error term representing the magnetic flux target, This represents the torque output value of the speed loop PI controller. Indicates the stator flux linkage reference value; Step 3-2: Determine priorities; First, normalize the error terms of the torque target and the flux linkage target. The normalization formula is as follows: , In the formula, This represents the normalized error of the torque. This represents the normalized error of the magnetic flux linkage. express These correspond to seven voltage vectors respectively; Then, the entropy values of the error terms for the torque target and the flux linkage target are calculated using the information entropy formula, which is: , In the formula, and The entropy values of the error terms for the torque target and the flux linkage target are respectively represented; then, the target with the larger entropy value is selected as the high priority, and the other target is selected as the low priority. Step 3-3: Target sorting; Assign error item ranking values to each error item of the torque target from smallest to largest; assign error item ranking values to each error item of the flux linkage target from smallest to largest. Steps 3-4: Determine the globally optimal vector; The cost function is decomposed to obtain the error terms of the torque target and the flux linkage target. Then, the eight vectors are traversed to obtain each error term of the torque target and the flux linkage target. Then, the priority of the torque target and the flux linkage target is dynamically determined according to the entropy value of each error term in the torque target and the flux linkage target, and they are respectively assigned as high priority and low priority. Then, the pre-selected vectors that meet the target tolerance threshold in the high-priority set are selected as the shortlisted vectors and added to the high-priority set. If the number of shortlisted vectors in the high-priority set is no more than one, and there is only one shortlisted vector, then the shortlisted vector is taken as the globally optimal vector. If there are zero shortlisted vectors, then the vector that minimizes the error term in the high-priority set among the eight pre-selected vectors is selected as the globally optimal vector. If the number of shortlisted vectors in the high-priority set is greater than one, then the vectors that meet another target tolerance threshold in the high-priority set are selected as the globally optimal vector. The shortlisted vectors are used as candidate vectors and enter the set of secondary priority vectors. If the number of candidate vectors in the set of secondary priority vectors is no more than 1, and there is 1 candidate vector, then the candidate vector is used as the global optimal vector. If there are 0 candidate vectors, then the vector that minimizes the error term in the secondary priority vector among the shortlisted vectors is selected as the global optimal vector. If the number of candidate vectors in the set of secondary priority vectors is greater than 1, then the vector with the smallest sum of the error term ranking values in the high priority set and the secondary priority set is selected as the global optimal vector.
7. The multi-objective model predictive control method based on tolerance-order optimization according to claim 1 or 6, characterized in that: In steps 3-4, the two target tolerance thresholds correspond to the torque tolerance threshold and the flux tolerance threshold. The two target tolerance thresholds are adaptively adjusted using the following algorithm, which specifically includes the following steps; Step S1: Establish the Actor network policy function; The correspondence between the operating state of the electric drive system and the output tolerance threshold is represented by an Actor network strategy function, which is: In the formula, This represents a function that retrieves the action. Represents an Actor network. Indicates the parameters of the Actor network; This represents the input motor environment state vector, where, In the formula, These represent the actual torque error and flux linkage error, respectively. Indicates the actual rotational speed; Step S2: Establish and update the mean square error loss function of the Critic network; wherein the mean square error loss function of the Critic network is: In the formula, For small batch sample size, This represents the loss function value of the Critic network. Indicates the parameters of the Critic network; Represents the Critic network, where, ; Indicates the first The target action value corresponding to the sample, obtained from the sampling of the first sample. The sample is The objective value function is: ; In the formula, Indicates the discount factor. Represents the target Actor network; Represents the target Critic network. and These represent the parameters of the target Actor network and the target Critic network, respectively. Indicates the first The current state of each sample; Indicates the execution of an action The environment then transitions to the state of the next moment; Indicates the first The action performed in the current state of a sample is usually output by the Actor network at that time; Indicates the execution of an action The immediate reward value returned by the environment is then used to evaluate the merits of the action. Step S3: The Actor network continuously optimizes the output policy using a policy gradient update method; The formula for the policy gradient update is: , ; In the formula, Indicates the parameters of the Actor network. The gradient; The action value function is expressed with respect to the action variable. The gradient; This represents the continuous actions output by the Actor network based on the current state; This represents the objective function for policy performance. Step S4: The Actor network parameters and the Critic network parameters are updated using gradient ascent and gradient descent methods, respectively, with the following formulas: , ; In the formula, and These are the learning rates for the Actor network and the Critic network, respectively. Step S5: Update the data from step S4. Substituting this into the Actor network policy function in step S1, we obtain the optimized tolerance threshold.
8. The multi-objective model predictive control method based on tolerance-order optimization according to claim 7, characterized in that: In step S4, the target network uses a soft update method, specifically, , ; In the formula, Denotes the target network soft update coefficient, and .
9. The multi-objective model predictive control method based on tolerance-order optimization according to claim 7, characterized in that: Before step S1, an experienced replay pool is established, defined as: In the formula, This indicates the system state at the next moment after the current tolerance threshold tuning action is performed.