Processor for controlling motor, motor control device, and motor control method

By introducing reinforcement learning calculator and PDFF controller in the motor control system, the problem of overshoot and parameter training time in the PID controller is solved, and more efficient motor handling performance and lower tracking errors are achieved.

CN120049773APending Publication Date: 2025-05-27IND TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410683857.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-05-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the existing motor control technology, the PID controller has problems with overshoot and long parameter training, and the motor speed and current tracking error is large.

Method used

The reinforcement learning calculator and pseudo-differential feedback and feedforward gain (PDFF) controller are used to improve overshoot and parameter training efficiency by applying reinforcement learning algorithms in the current loop of the PID controller and using PDFF controllers in the speed loop, and adjust the instantaneous response speed through the feedforward proportional coefficient.

Benefits of technology

It effectively reduces the tracking error of rotation speed and current in the motor, improves the handling performance, and reduces the time-consuming parameter tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120049773A_ABST
    Figure CN120049773A_ABST
Patent Text Reader

Abstract

The invention provides a processor for controlling a motor, a motor control device and a control method. The processor includes a feedback calculator, a control calculator, and a drive calculator. The feedback calculator calculates a direct-axis current and a quadrature-axis current on the basis of a drive current for driving the motor and an operating angle of the motor. The control calculator includes a reinforcement learning controller. The reinforcement learning controller uses a reinforcement learning algorithm to calculate a direct-axis voltage and a quadrature-axis voltage according to the quadrature-axis current command, the direct-axis current and the quadrature-axis current. The quadrature-axis current command is obtained according to the reference rotating speed and the running speed of the motor. The driving calculator generates a switching signal according to the direct-axis voltage, the quadrature-axis voltage and the operation angle. The switching signal is used for controlling the driving circuit to drive the motor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a processor for controlling an electric motor, a motor control device, and a control method for an electric motor. Background Art

[0002] Current means of transportation are mainly developed in the direction of electric vehicles or electric drive-assisted vehicles, and related technologies such as electric drive-assisted vehicles have diverse applications. The most important aspects of electric vehicles are power supply and electric motor drive.

[0003] Electric motor drive technology often uses field-oriented control technology and a proportional-integral-derivative (PID) controller to achieve the drive and control of an electric motor. However, since electric vehicles often face unexpected dynamic changes in torque load, rotor resistance, or stator resistance, and different specifications of electric vehicle motors and different degrees of torque load changes require adjusting the parameters in the PID controller one by one to optimize the drive control performance of the motor. Therefore, how to improve field-oriented control technology and effectively enhance the control performance of electric motors is one of the research directions. Summary of the Invention

[0004] The present invention provides a processor for controlling an electric motor, a motor control device, and a control method, which can improve the overshoot problem in a proportional-integral-derivative (PID) controller and the time-consuming situation of parameter adjustment, and reduce the tracking error between the rotational speed and current in the motor.

[0005] According to an embodiment of the present invention, a processor for controlling an electric motor includes a feedback calculator, a control calculator, and a drive calculator. The feedback calculator calculates a direct-axis current and a quadrature-axis current based on a drive current for driving the electric motor and a rotation angle of the electric motor. The control calculator is coupled to the feedback calculator. The control calculator includes a reinforcement learning controller. The reinforcement learning controller uses a reinforcement learning algorithm to calculate a direct-axis voltage and a quadrature-axis voltage based on a quadrature-axis current command, the direct-axis current, and the quadrature-axis current. The quadrature-axis current command is obtained based on a reference rotational speed and the rotational speed of the electric motor. The drive calculator is coupled to the control calculator. The drive calculator generates a switching signal based on the direct-axis voltage, the quadrature-axis voltage, and the rotation angle. The switching signal is used to control a drive circuit to drive the electric motor.

[0006] According to an embodiment of the present invention, a motor control device includes a processor, a drive circuit, and a sensor. The drive circuit is coupled to the processor and is controlled by the processor to drive a motor. The sensor is coupled to the processor. The sensor is used to sense the operating speed and operating angle of the motor. The processor controls the drive circuit based on the drive current of the drive circuit, the operating speed, and the operating angle of the motor. The processor includes a feedback calculator, a control calculator, and a drive calculator. The feedback calculator calculates a direct-axis current and a quadrature-axis current based on the drive current and the operating angle of the motor. The control calculator is coupled to the feedback calculator. The control calculator includes a reinforcement learning controller. The reinforcement learning controller uses a reinforcement learning algorithm to calculate a direct-axis voltage and a quadrature-axis voltage based on a quadrature-axis current command, the direct-axis current, and the quadrature-axis current. The quadrature-axis current command is obtained based on a reference speed and the operating speed of the motor. The drive calculator is coupled to the control calculator. The drive calculator generates a switching signal based on the direct-axis voltage, the quadrature-axis voltage, and the operating angle. The switching signal is used to control the drive circuit to drive the motor.

[0007] According to an embodiment of the present invention, a control method for a motor includes the following steps: sensing the operating speed and operating angle of the motor; calculating a direct-axis current and a quadrature-axis current based on a drive current for driving the motor and the operating angle; using a reinforcement learning algorithm to calculate a direct-axis voltage and a quadrature-axis voltage based on a quadrature-axis current command, the direct-axis current, and the quadrature-axis current, wherein the quadrature-axis current command is obtained based on a reference speed and the operating speed of the motor; and generating a switching signal based on the direct-axis voltage, the quadrature-axis voltage, and the operating angle, wherein the switching signal is used to control a drive circuit to drive the motor.

[0008] Based on the above, in the current loop of the PID controller, the processor, the motor control device, and the control method for controlling a motor according to the embodiments of the present invention adopt a reinforcement learning calculator and a reinforcement learning algorithm applied to motor control, and in the speed loop of the PID controller, a PDFF controller in the control calculator is used to improve the overshoot problem in the PID controller and the time-consuming situation of parameter tuning, and the instantaneous response speed is adjusted by the feedforward proportional coefficient in the PDFF controller to reduce the tracking error between the speed and the current in the motor. Therefore, the control performance of the controlled motor can be effectively improved. Description of the Drawings

[0009] Figure 1 is a schematic diagram of a motor control device according to a first embodiment of the present invention.

[0010] Figure 2A and 2BIt is a schematic diagram of implementing a reinforcement learning algorithm using a reinforcement learning controller in the first embodiment of the present invention.

[0011] Figure 3 It is a schematic diagram of a motor control device according to the second embodiment of the present invention.

[0012] Figure 4 It is Figure 3 a schematic diagram in which the PDFF controller is used to calculate the quadrature axis current command.

[0013] Figure 5 It is Figure 1 a schematic diagram of comparing the current loop performance between a processor and a PID controller implemented by a PI controller in the first embodiment.

[0014] Figure 6 It is Figure 1 a schematic diagram of comparing the speed loop performance between a processor and a PID controller implemented by a PI controller in the first embodiment.

[0015] Figure 7 It is Figure 3 a schematic diagram of comparing the speed loop performance between a processor and a PID controller implemented by a PI controller in the second embodiment.

[0016] Figure 8 It is a flowchart of a control method for a motor according to an embodiment of the present invention. Detailed implementation manners

[0017] Now, reference will be made in detail to the exemplary embodiments of the present invention, and examples of the exemplary embodiments are illustrated in the accompanying drawings. Whenever possible, the same component symbols are used in the drawings and the description to represent the same or similar parts.

[0018] Proportional-Integral-Derivative (PID) controllers often use multiple Proportional-Integral (PI) controllers to implement the current loop and speed loop in the PID controller. However, large overshoots often occur in the voltage commands generated by the PID controller, and its adaptability to the overall system parameters and external disturbances in the motor control device is poor. The "current loop" means that the PID controller sets the magnitude of the output torque of the motor shaft to the outside through the input of external data or through simulation, and is applied to situations where the torque of the motor needs to be strictly controlled, so as to be used as the control of the current loop. The "speed loop" means that the PID controller controls the rotational speed of the motor through the input of external data or through simulation.

[0019] In the embodiment of the present invention, a reinforcement learning calculator is adopted in the current loop of a proportional-integral-derivative (PID) controller and a reinforcement learning algorithm applied to motor control, and a pseudo-differential feedback and feedforward gain (PDFF) controller in the control calculator is utilized in the speed loop of the PID controller to improve the overshoot problem in the PID controller and the time-consuming situation of parameter tuning, and enhance the control performance of the controlled motor. Multiple embodiments are proposed below for further illustration.

[0020] Figure 1 FIG. 4 is a schematic diagram of a motor control device 100 according to a first embodiment of the present invention. The motor control device 100 is used to drive a motor 105. In this embodiment, the motor 105 is exemplified by a permanent-magnet synchronous motor (PMSM). The motor control device 100 mainly includes a processor 110, a drive circuit 120, and a sensor 130.

[0021] The processor 110 can be implemented by a logic circuit. For example, the processor 110 can be a microprocessor. The drive circuit 120 is coupled to the processor 110 and the motor 105. The drive circuit 120 is controlled by the processor 110 to drive the motor 105. The sensor 130 is coupled to the processor 110 and the motor 105. The sensor 130 senses the operating speed ω and the operating angle θ of the motor 105, and provides the operating speed ω and the operating angle θ to the processor 110. The operating speed ω is the rotational speed of the motor, and its unit can be revolutions per minute (RPM). The processor 110 generates a switching signal SWS based on the drive current of the drive circuit 120 (such as, Figure 1 the drive currents ia and ib in FIG. 4), the operating speed ω and the operating angle θ of the motor 105, and controls the drive circuit 120 through the switching signal SWS. The drive circuit 120 generates a corresponding drive current according to the switching signal to drive the motor 105.

[0022] The processor 110 mainly includes a control calculator 111, a drive calculator 114, and a feedback calculator 116. The feedback calculator 116 calculates the direct-axis current id and the quadrature-axis current iq by performing a coordinate transformation on the current based on the drive current (such as, Figure 1 the drive currents ia and ib in FIG. 4) for driving the motor 105 and the operating angle θ of the motor 105.

[0023] Specifically, the feedback calculator 116 includes a Clarke transformation controller 117-1 and a Park transformation controller 117-2. The Clarke transformation controller 117-1 transforms the drive current located in the time-domain coordinate system (such as, Figure 1The drive currents ia and ib are transformed into a first current iα and a second current iβ in an orthogonal stationary coordinate system (denoted by αβ). The Park transformation controller 117-2 is coupled to the Clarke transformation controller 117-1. The Park transformation controller 117-2 transforms the first current iα and the second current iβ in the orthogonal stationary coordinate system (denoted by αβ) into a direct-axis current id and a quadrature-axis current iq in an orthogonal rotating coordinate system (denoted by dq).

[0024] The control calculator 111 is coupled to the feedback calculator 116. The control calculator 111 may include a reinforcement learning controller 112 and a proportional-integral (PI) controller 113. The reinforcement learning controller 112 uses the reinforcement learning algorithm of the embodiment of the present invention to calculate a direct-axis voltage Vd and a quadrature-axis voltage Vq based on the quadrature-axis current command iqref, the direct-axis current id, and the quadrature-axis current iq. Details related to the reinforcement learning controller 112 and the reinforcement learning algorithm are described below Figure 2A 、 2B and corresponding descriptions.

[0025] The quadrature-axis current command iqref of this embodiment is obtained based on a reference rotational speed Wref and the operating speed W of the motor 105. Specifically, the first embodiment of the present invention uses the PI controller 113 and the subtractor 118 to generate the quadrature-axis current command iqref based on the difference between the operating speed W and the reference rotational speed Wref. Those applying this embodiment may also use other methods to generate the quadrature-axis current command iqref, as long as the quadrature-axis current command iqref is obtained based on the reference rotational speed Wref and the operating speed W of the motor 105.

[0026] The drive calculator 114 is coupled to the control calculator 111. The drive calculator 114 generates a switching signal SWS based on the direct-axis current id, the quadrature-axis current iq, and the operating angle θ. The switching signal SWS is used to control the drive circuit 120 to drive the motor 105. Specifically, the drive calculator 114 includes a Park inverse transformation controller 115-1 and a Clarke inverse transformation controller 115-2. The Park inverse transformation controller 115-1 transforms the direct-axis voltage Vd and the quadrature-axis voltage Vq in the orthogonal rotating coordinate system dq into a first voltage Vα and a second voltage Vβ in the orthogonal stationary coordinate system αβ. The Clarke inverse transformation controller 115-2 is coupled to the Park inverse transformation controller 115-1. The Clarke inverse transformation controller 115-2 transforms the first voltage Vα and the second voltage Vβ in the orthogonal stationary coordinate system αβ into the switching signal SWS.

[0027] The processor 110 further includes a subtractor 118 and a zero current supplier 119. The subtractor 118 subtracts the operating speed W and the reference rotational speed Wref from each other to generate a difference therebetween, and supplies this difference to the PI controller 113. The zero current supplier 119 is coupled to the reinforcement learning controller 112. The zero current supplier 119 is configured to supply a zero current as the direct-axis current command idref. The reinforcement learning controller 112 can use a reinforcement learning algorithm to calculate the direct-axis voltage Vd and the quadrature-axis voltage Vq according to the quadrature-axis current command iqref, the direct-axis current command idref, the direct-axis current id, and the quadrature-axis current iq. In this embodiment, the direct-axis current command idref is set to be the zero current supplied by the zero current supplier 119.

[0028] Figure 2A and 2B is a schematic diagram of implementing a reinforcement learning algorithm by using the reinforcement learning controller 112 in the first embodiment of the present invention. Figure 2A A schematic diagram showing the relationship between the environment 210, the observation items 220, the action items 240, the decision 230, and the reinforcement learning algorithm 205. The reinforcement learning algorithm 205 is an effective method for solving sequential decision-making problems. The reinforcement learning algorithm 205 can also be referred to as an agent. The environment 210 is the world that interacts with the agent. In each step of the interaction, the agent obtains the observation items 220 of the state of the environment 210 it is in, and then relies on the decision 230 to determine the action to be executed in the next step. The environment 210 will change due to the actions of the agent on it, or it may change by itself. The agent also perceives a current reward 250 from the environment that indicates the quality of the current state. The goal of the agent is to maximize the cumulative current reward.

[0029] As Figure 2A shown, the reinforcement learning algorithm 205 mainly includes a decision 230 and a reinforcement learning control training algorithm 260. The decision 230 is an equation self-adjusted by the reinforcement learning algorithm 205, so the decision 230 can also be referred to as a decision equation. The reinforcement learning control training algorithm 260 is an operation logic algorithm and corresponding technology for adjusting the decision equation. The reinforcement learning algorithm 205 is a technology by which the agent continuously corrects its own decision 230 through learning behaviors to achieve the goal.

[0030] In this embodiment, under environment 210, the following four values are mainly observed as observation items 220: direct-axis current id, quadrature-axis current iq, direct-axis current error value iderror calculated from the difference between the current direct-axis current id and the previous direct-axis current, and quadrature-axis current error value iqerror calculated from the difference between the current quadrature-axis current iq and the previous quadrature-axis current. The direct-axis voltage Vd and the quadrature-axis voltage Vq are used as action items 240 of the reinforcement learning algorithm.

[0031] The input of the reinforcement learning algorithm 205 is mainly each value in the observation item 220, and the output of the reinforcement learning algorithm 205 is each value in the action item 240. The decision 230 in the reinforcement learning algorithm 205 mainly uses each value in the observation item 220 for calculation and transforms it into each value in the action item 240. The reinforcement learning control training algorithm 260 in the reinforcement learning algorithm 205 determines whether to perform decision update 235 and determines the adjustment degree for the decision update 235 according to the current reward 250.

[0032] Figure 2A Based on the reward equation, the current reward 250 is calculated according to the corresponding data of the current observation item 220 and the current action item 240. Specifically, the current reward rt 250 can be calculated by the following reward equation (1):

[0033]

[0034] "iderror" in the reward equation (1) is the aforementioned direct-axis current error value, "iqerror" is the aforementioned quadrature-axis current error value, Q1, Q2, and R are preset parameters, and "rt" is the current reward 250. "j" represents the action index. is the action at the previous time step. In this embodiment, Q1 and Q2 are set to 5, and R is set to 0.1. Those who apply this embodiment can adjust the preset parameters such as Q1, Q2, and R according to their needs.

[0035] Figure 2B It is a schematic diagram of the reinforcement learning algorithm presented by multiple functional blocks using simulation software (such as MATLAB / Simulink). Figure 2BThe observation items 220 include the direct-axis current id, the quadrature-axis current iq, the direct-axis current error value iderror, and the quadrature-axis current error value iqerror. The current reward 250 is mainly calculated from the direct-axis current error value iderror, the quadrature-axis current error value iqerror, and the previous action item 240. The reinforcement learning algorithm 205 calculates the predicted action item 240 from the aforementioned observation items 220, the current reward 250, and the completed data (e.g., the zero current provided by the zero current supplier 119 as the direct-axis current command idref). The reinforcement learning algorithm 205 of this embodiment can also be referred to as a twin delayed deep deterministic policy gradients (TD3) agent.

[0036] Those applying this embodiment can use different types of reinforcement learning algorithms according to their needs to implement Figure 1 the reinforcement learning controller 112. Here is an example to illustrate Figure 2A the training steps of the reinforcement learning control training algorithm 260. The training steps of the reinforcement learning control training algorithm 260 can be mainly divided into Step 1 to Step 6.

[0037] In Step 1, a specific action item is selected. In this embodiment, action A is selected and presented by the following equation (2):

[0038] A = μ(S) + N…(2)

[0039] In equation (2) corresponding to action A, “S” is the current state, and “N” is random noise.

[0040] After selecting the specific action item (i.e., action A), the second step (Step 2) is executed. Step 2 includes the following sub-steps 1 to sub-step 3. Sub-step 1 is to execute the selected action A to generate an action value AV. Sub-step 2 is to calculate the aforementioned current reward rt based on the aforementioned reward equation (1). Sub-step 3 is to calculate the corresponding state of the next observation item as the state data S’. After executing sub-steps 1 to 3, the current state S, the action value AV, the current reward rt, and the state data S’ are stored as a set of training patterns, which are presented as a set of training patterns (S, AV, rt, S’).

[0041] In Step 3, the aforementioned Step 2 is executed multiple times (e.g., executed M times, where M is a positive integer) to randomly generate multiple sets of training patterns.

[0042] Step 4 is to calculate multiple value function targets yi based on these multiple sets of training patterns. The equation (3) of the value function target yi is presented as follows:

[0043] yᵢ = Rᵢ + γ·min(Qₖ′(Sₖ′, clip(μ′(Sₖ′|θᵤ) + ε)|θ Qk′ ))…(3)

[0044] In Equation (3), "Rᵢ" is the reward, and the value function target yᵢ is the sum of the reward Rᵢ and the minimum discounted future reward of the critics. "Qₖ′" is the action value function for policy k. "Sₖ′" is the state for policy k. "θᵤ" is the parameter representing the asynchronous work item. "θ Qk′ " is the action value function in the asynchronous work item.

[0045] Step 5 is to update each critics parameter to minimize the parameter Lₖ. The equation for the parameter Lₖ is presented as follows:

[0046]

[0047] In Equation (3), "Qₖ" is the action value function for policy k, "Sᵢ" is the state, and "Aᵢ" is the action. "θ Qk " is the action value function in the asynchronous work item.

[0048] Step 6 is to update the parameter in action A to maximize the reward. The equation for maximizing the reward is presented as follows:

[0049]

[0050] The parameter G in Equation (5) ai The corresponding Equation (6) is presented as follows:

[0051]

[0052] The parameter G in Equation (5) ui The corresponding Equation (7) is presented as follows:

[0053]

[0054] The corresponding Equation (8) for the parameter A in Equation (6) is presented as follows:

[0055] A = μ(S i |θ μ )…(8)

[0056] After performing Steps 1 to 6, Figure 2A the reinforcement learning control training algorithm 260 in can correspondingly adjust the equations in the decision 230 through the decision update 235, thereby realizing the function of the deep neural network.

[0057] Figure 3 It is a schematic diagram of a motor control device 300 according to the second embodiment of the present invention. Figure 1 Compared with Figure 3 The main difference is that the second embodiment uses the pseudo-differential feedback and feed-forward gain (PDFF) controller 313 and the subtractor 118 in the processor 310 to generate the quadrature-axis current command iqref based on both the operating speed W and the reference speed Wref. Specifically, the PDFF controller 313 calculates the quadrature-axis current command iqref according to the reference speed Wref and the operating speed W of the motor.

[0058] Figure 4 is Figure 3 A schematic diagram of the PDFF controller 313 in calculating the quadrature-axis current command iqref. The PDFF controller 313 calculates the quadrature-axis current command iqref according to the following equation (9):

[0059]

[0060] "W" is the operating speed of the motor, "Wref" is the preset reference speed in this embodiment, "r" is the feed-forward proportional coefficient, "Kpf" is the feedback proportional gain, "KI" is the integral gain, is the Z-transform value of the integral gain, and "iqref" is the quadrature-axis current command.

[0061] The equation (9) in the PDFF controller 313 is applied to the processor 310 (for example, a PID controller) in the default formula form, and the foregoing equation (9) does not require training. Therefore, this embodiment adopts the quadrature-axis equivalent stator current command (such as the quadrature-axis current command iqref) output by the PDFF controller 313 in the speed loop of the PID controller, so that it can effectively eliminate the overshoot and can adjust the instantaneous response speed through the foregoing various gains and coefficients (such as the feed-forward proportional coefficient r, the feedback proportional gain Kpf, the integral gain KI, etc.), and reduce the tracking error of the input data.

[0062] Figure 5 is Figure 1 A schematic diagram of the comparison of the current loop performance between the processor 110 in the first embodiment and the PID controller implemented by using a PI controller. Figure 5 In which, the horizontal axis is time and the vertical axis is the measured error of the quadrature-axis current value (in amperes). From Figure 5 it can be seen that Figure 1The waveform 510 (represented by a solid line) of the quadrature-axis current value error of the processor 110 using the reinforcement learning controller 112 fluctuates significantly less than the corresponding waveform 520 (represented by a dashed line) of the quadrature-axis current value error generated by the PID controller implemented using a PI controller. After simulation, Figure 1 The processor 110 using the reinforcement learning controller 112 in [the relevant context] can reduce the quadrature-axis current value error by greater than or equal to 30%.

[0063] Figure 6 is Figure 1 A schematic diagram comparing the speed loop performance between the processor 110 and the PID controller implemented using a PI controller in the first embodiment. Figure 6 In [the relevant context], the horizontal axis is time and the vertical axis is the measured operating speed (in revolutions per minute (RPM)). From Figure 6 it can be seen that Figure 1 The waveform 610 (represented by a solid line) of the operating speed between the processor 110 using the reinforcement learning controller 112 and the estimated speed fluctuates significantly less than the corresponding waveform 620 (represented by a dashed line) of the operating speed and the estimated speed of the PID controller implemented using a PI controller. After simulation, Figure 1 The processor 111 using the reinforcement learning controller 112 in [the relevant context] can reduce the speed error by greater than or equal to 10%.

[0064] Figure 7 is Figure 3 A schematic diagram comparing the speed loop performance between the processor 311 and the PID controller implemented using a PI controller in the second embodiment. Figure 7 In [the relevant context], the horizontal axis is time and the vertical axis is the measured operating speed (in revolutions per minute (RPM)). From Figure 7 it can be seen that Figure 3 The waveform 710 (represented by a solid line) of the operating speed between the processor 310 using the PDFF controller 313 and the reinforcement learning controller 112 and the estimated speed fluctuates significantly less than the corresponding waveform 720 (represented by a dashed line) of the operating speed and the estimated speed of the PID controller implemented using a PI controller. After simulation, Figure 3 The processor 310 using the PDFF controller 313 and the reinforcement learning controller 112 in [the relevant context] can reduce the speed error by greater than or equal to 30%.

[0065] Figure 8 It is a flowchart of a control method for a motor according to an embodiment of the present invention. Figure 8 The control method can be applied to Figure 1 the corresponding hardware structure of the first embodiment or Figure 3 the corresponding hardware structure of the second embodiment. Here, taking Figure 1 in combination with Figure 8 to illustrateFigure 8 Control method. In step S810, use Figure 1 sensor 130 to sense the operating speed W and operating angle θ of the motor 105. In step S820, use Figure 1 processor 110 to calculate the direct-axis current id and quadrature-axis current iq based on the drive current (e.g., Figure 1 drive currents ia, ib) for driving the motor 105 and the operating angle θ.

[0066] In step S830, use Figure 1 the reinforcement learning controller 112 in the processor 110 and use the reinforcement learning algorithm to calculate the direct-axis voltage Vd and quadrature-axis voltage Vq based on the quadrature-axis current command iqref, direct-axis current id, and quadrature-axis current iq. The quadrature-axis current command iqref is obtained by using Figure 1 the PI controller 113 in Figure 3 the PDFF controller 313 in Figure 1 according to the reference speed Wref and the operating speed W of the motor 105. In step S840, use

[0067] Figure 8 For the detailed processes of steps S810 to S840 of the control method in

[0068] See the foregoing embodiments.

[0069] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A processor for controlling an electric motor, characterized in that: include: A feedback calculator for calculating a direct-axis current and a quadrature-axis current according to a driving current for driving the motor and a running angle of the motor; a control calculator coupled to the feedback calculator, the control calculator comprising a reinforcement learning controller, wherein the reinforcement learning controller uses a reinforcement learning algorithm to calculate a direct-axis voltage and a quadrature-axis voltage according to a quadrature-axis current command, the direct-axis current, and the quadrature-axis current, wherein the quadrature-axis current command is obtained according to a reference speed and a running speed of the motor; as well as A driving calculator is coupled to the control calculator and generates a switch signal according to the direct-axis voltage, the quadrature-axis voltage and the operating angle. The switch signal is used to control a driving circuit to drive the motor.

2. The processor according to claim 1, characterized in that The reinforcement learning algorithm uses the direct-axis current, the cross-axis current, the direct-axis current error value and the cross-axis current error value as observation items of the reinforcement learning algorithm, uses the previous direct-axis voltage and the cross-axis voltage as action items of the reinforcement learning algorithm, calculates the current reward based on the corresponding data of the observation items and the action items based on the reward equation, and calculates the estimated action item based on the observation items, the current reward and the completed data based on the decision equation in the reinforcement learning algorithm and the reinforcement learning control training algorithm, wherein the estimated action item includes the direct-axis voltage and the cross-axis voltage.

3. The processor according to claim 2, characterized in that The reward equation is: Among them, iderror is the direct-axis current error value, iqerror is the quadrature-axis current error value, Q1, Q2 and R are preset parameters, and rt is the current reward.

4. The processor according to claim 1, characterized in that The training steps of the reinforcement learning control training algorithm include: Selecting a first action, the first action comprising a current state and random noise; executing a second step, the second step comprising executing the first action to generate an action value, calculating the current reward based on the reward equation, calculating a corresponding state of the next observation item as state data, and storing the current state, the action value, the current reward, and the state data as a set of training patterns; Performing the second step multiple times to randomly generate multiple sets of training patterns; Calculating a plurality of value function targets based on the plurality of sets of training patterns; and The review parameters in the neural network are corrected based on the multiple sets of training patterns and the multiple value function targets to train the reinforcement learning control training algorithm.

5. The processor according to claim 1, wherein: The control calculator also includes: A pseudo differential feedback and feedforward gain controller is coupled to the reinforcement learning controller and calculates the quadrature-axis current command according to the reference speed and the running speed of the motor.

6. The processor according to claim 1, wherein: The pseudo-differential feedback and feedforward gain controller calculates the quadrature-axis current command according to the following equation: Wherein, W is the running speed of the motor, Wref is the reference speed, r is the feedforward proportional coefficient, Kpf is the feedback proportional gain, KI is the integral gain, Z -1 is the Z transform, iqref is the quadrature-axis current command.

7. The processor according to claim 1, wherein: The reinforcement learning algorithm is a twin delayed deep deterministic policy gradient algorithm.

8. The processor according to claim 1, wherein: The feedback calculator comprises: A Clarke transform controller transforms the driving current in a time domain coordinate system into a first current and a second current in an orthogonal stationary coordinate system; and The Park transform controller is coupled to the Clarke transform controller and transforms the first current and the second current in the orthogonal stationary coordinate system into the direct-axis current and the quadrature-axis current in an orthogonal rotating coordinate system.

9. The processor according to claim 7, characterized in that: The drive calculator comprises: an inverse Park transform controller, which transforms the direct-axis voltage and the quadrature-axis voltage in the orthogonal rotating coordinate system into a first voltage and a second voltage in the orthogonal stationary coordinate system; and The Clarke transform controller is coupled to the inverse Park transform controller and transforms the first voltage and the second voltage in an orthogonal stationary coordinate system into the switching signal.

10. The processor according to claim 1, wherein: Also includes: a zero current supplier, coupled to the reinforcement learning controller, for providing zero current as a direct-axis current command, The reinforcement learning controller utilizes the reinforcement learning algorithm to calculate the direct-axis voltage and the quadrature-axis voltage according to the quadrature-axis current command, the direct-axis current command, the direct-axis current, and the quadrature-axis current.

11. A motor control device, characterized in that: include: processor; a driving circuit, coupled to the processor and controlled by the processor to drive the motor; as well as A sensor, coupled to the processor, for sensing the running speed and running angle of the motor, wherein the processor controls the drive circuit according to the drive current of the drive circuit, the running speed of the motor and the running angle, The processor comprises: A feedback calculator, for calculating a direct-axis current and a quadrature-axis current according to the driving current and the running angle of the motor; a control calculator coupled to the feedback calculator, the control calculator comprising a reinforcement learning controller, wherein the reinforcement learning controller uses a reinforcement learning algorithm to calculate a direct-axis voltage and a quadrature-axis voltage according to a quadrature-axis current command, the direct-axis current, and the quadrature-axis current, wherein the quadrature-axis current command is obtained according to a reference speed and the operating speed of the motor; and The driving calculator is coupled to the control calculator and generates a switch signal according to the direct-axis voltage, the quadrature-axis voltage and the operating angle, wherein the switch signal is used to control the driving circuit.

12. The motor control device according to claim 11, characterized in that: The reinforcement learning controller uses the reinforcement learning algorithm to calculate the direct-axis voltage and the quadrature-axis voltage, including: Using the direct-axis current, the quadrature-axis current, the direct-axis current error value, and the quadrature-axis current error value as observation items of the reinforcement learning algorithm; Using the previous direct-axis voltage and the quadrature-axis voltage as action items of the reinforcement learning algorithm; Calculating a current reward based on a reward equation according to corresponding data of the observation item and the action item; and Based on a reinforcement learning control training algorithm, an estimated action item is calculated according to the observation item, the current reward and the completed data, wherein the estimated action item includes the direct-axis voltage and the quadrature-axis voltage.

13. The motor control device according to claim 12, characterized in that: The reward equation is: Among them, iderror is the direct-axis current error value, iqerror is the quadrature-axis current error value, Q1, Q2 and R are preset parameters, and rt is the current reward.

14. The motor control device according to claim 13, characterized in that: The training steps of the reinforcement learning control training algorithm include: Selecting a first action, the first action comprising a current state and random noise; executing a second step, the second step comprising generating an action value by the first action, calculating the reward based on the reward equation, calculating a corresponding state of the next observation item as state data, and storing the current state, the action value, the reward, and the state data as a set of training patterns; Performing the second step multiple times to randomly generate multiple sets of training patterns; Calculating a plurality of value function targets based on the plurality of sets of training patterns; and The review parameters in the neural network are corrected based on the multiple sets of training patterns and the multiple value function targets to train the reinforcement learning control training algorithm.

15. The motor control device according to claim 11, characterized in that: The control calculator also includes: A pseudo differential feedback and feedforward gain controller is coupled to the reinforcement learning controller and calculates the quadrature-axis current command according to the reference speed and the running speed of the motor.

16. The motor control device according to claim 15, characterized in that: The pseudo-differential feedback and feedforward gain controller calculates the quadrature-axis current command according to the following equation: Wherein, W is the running speed of the motor, W is the reference speed, r is the feedforward proportional coefficient, Kpf is the feedback proportional plus one, KI is the integral gain, and Z -1 is the Z transform, iqref is the quadrature-axis current command.

17. The motor control device according to claim 11, characterized in that: The reinforcement learning algorithm is a twin delayed deep deterministic policy gradient algorithm.

18. A control method for an electric motor, characterized in that: include: sensing the running speed and running angle of the motor; Calculating a direct-axis current and a quadrature-axis current according to a driving current for driving the motor and the operating angle; Utilizing a reinforcement learning algorithm to calculate a direct-axis voltage and a quadrature-axis voltage according to a quadrature-axis current command, the direct-axis current, and the quadrature-axis current, wherein the quadrature-axis current command is obtained according to a reference speed and the operating speed of the motor; as well as A switch signal is generated according to the direct-axis voltage, the quadrature-axis voltage and the operating angle, wherein the switch signal is used to control a driving circuit to drive the motor.

19. The control method according to claim 18, characterized in that: The step of using a reinforcement learning algorithm to calculate the direct-axis voltage and the quadrature-axis voltage according to the quadrature-axis current command, the direct-axis current and the quadrature-axis current comprises: Using the direct-axis current, the quadrature-axis current, the direct-axis current error value, and the quadrature-axis current error value as observation items of the reinforcement learning algorithm; Using the previous direct-axis voltage and the quadrature-axis voltage as action items of the reinforcement learning algorithm; Calculating a current reward based on a reward equation according to corresponding data of the observation item and the action item; and Based on a reinforcement learning control training algorithm, an estimated action item is calculated according to the observation item, the current reward and the completed data, wherein the estimated action item includes the direct-axis voltage and the quadrature-axis voltage.

20. The control method according to claim 19, characterized in that: The training steps of the reinforcement learning control training algorithm include: Selecting a first action, the first action comprising a current state and random noise; executing a second step, the second step comprising generating an action value by the first action, calculating the reward based on the reward equation, calculating a corresponding state of the next observation item as state data, and storing the current state, the action value, the reward, and the state data as a set of training patterns; Performing the second step multiple times to randomly generate multiple sets of training patterns; Calculating a plurality of value function targets based on the plurality of sets of training patterns; and The review parameters in the neural network are corrected based on the multiple sets of training patterns and the multiple value function targets to train the reinforcement learning control training algorithm.