Motor and control method and device thereof, compressor, air conditioner, medium and product

By combining fuzzy logic and DQN reinforcement learning FOC control, the inverter control in motor drive control is optimized, and the problem of difficulty in setting PI controller parameters is solved, and the effect and system performance of motor drive control are improved.

CN120357791APending Publication Date: 2025-07-22GREE ELECTRIC APPLIANCE INC OF ZHUHAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510641787.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In motor drive control, the d-axis and q-axis control loops of the FOC control three-phase inverter require two PI controllers, which leads to difficulty in setting parameters under different power supply and load conditions, affecting the control effect.

Method used

The FOC control method combining fuzzy logic processing and DQN reinforcement learning is adopted to replace the PI controller of the d-axis and q-axis, and the PI controller of the speed ring and the magnetic flux ring are retained. By obtaining the three-phase current of the motor, the actual speed and the rotor magnetic flux target quantity, the fuzzy logic and the DQN reinforcement learning algorithm are used to optimize the switching tube control signal of the inverter module.

Benefits of technology

Improve the effect of motor drive control, avoid difficulty in parameter setting, and improve the energy efficiency, stability and performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357791A_ABST
    Figure CN120357791A_ABST
Patent Text Reader

Abstract

The invention discloses a motor and a control method and device thereof, a compressor, an air conditioner, a storage medium and a computer program product. The method comprises the following steps: acquiring a three-phase current of the motor, acquiring an actual rotating speed of the motor, acquiring a target rotating speed of the motor, and acquiring a rotor flux linkage target quantity of the motor; according to the three-phase current of the motor, the actual rotating speed of the motor, the target rotating speed of the motor and the rotor flux linkage target quantity of the motor, after FOC control combining fuzzy logic processing and DQN reinforcement learning processing is adopted, a control signal of a switching tube in an inversion module is obtained; and according to a control signal of a switch tube in the inversion module, the inversion module is controlled to work so as to realize control of the motor. According to the scheme, on the basis of the three-phase current and the rotating speed of the motor, the FOC control combining fuzzy logic processing and DQN reinforcement learning processing is adopted, the control signal of the switching tube in the inversion module is controlled to achieve driving control over the motor, and the control effect of the motor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of motors, and particularly relates to a control method, device, motor, compressor, air conditioner, storage medium and computer program product of a motor, and more particularly to a motor drive control method, device, motor, compressor, air conditioner, storage medium and computer program product for an air conditioner compressor based on field-oriented control of DQN reinforcement learning optimized by fuzzy logic. Background Art

[0002] Three-phase inverters are widely used in power electronic systems, especially in the fields of new energy power generation, motor drive and power transmission. In motor drive control (such as motor drive control in an air conditioner compressor), in related solutions, field-oriented control (FOC) is adopted for the control of the three-phase inverter, and a loop controller, such as PI (proportional-integral) control, needs to be set in a key loop. When PI control is used, there are two control loops, one d-axis control loop and one q-axis control loop, and two PI controllers are required for the d-axis control loop and the q-axis control loop; the two PI controllers have problems with difficult parameter tuning under various different power supplies and loads, which affects the control effect of the motor.

[0003] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] An object of the present invention is to provide a control method, device, motor, compressor, air conditioner, storage medium and computer program product of a motor, so as to solve the problem that in motor drive control (such as motor drive control in an air conditioner compressor), when FOC is used to control a three-phase inverter in related solutions, two PI controllers are required for the d-axis control loop and the q-axis control loop, and there are problems with difficult parameter tuning under different power supplies and loads, which affects the control effect of the motor, and achieve the effect of adopting FOC control combined with fuzzy logic processing and DQN reinforcement learning processing based on the three-phase current and speed of the motor to control the inverter to realize the drive control of the motor and improve the control effect of the motor.

[0005] The present invention provides a control method for a motor. The drive control end of the motor is provided with an inverter module, and the inverter module is provided with switching tubes. The control method for the motor includes: obtaining the three-phase current of the motor, obtaining the actual speed of the motor, obtaining the target speed of the motor, and obtaining the target value of the rotor magnetic flux of the motor; according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, after performing FOC control combining fuzzy logic processing and DQN reinforcement learning processing, obtaining the control signal of the switching tubes in the inverter module; according to the control signal of the switching tubes in the inverter module, controlling the operation of the inverter module to achieve the control of the motor.

[0006] In some embodiments, according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, after performing FOC control combining fuzzy logic processing and DQN reinforcement learning processing, obtaining the control signal of the switching tubes in the inverter module includes: according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, after being estimated by an estimator, transformed by a converter, and PI controlled by a PI controller, obtaining the weight coefficient of the reward function in the DQN reinforcement learning processing, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor; according to the weight coefficient of the reward function, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor, after performing DQN reinforcement learning processing by the DQN reinforcement learning algorithm, obtaining the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor; according to the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor, after performing inverse Park transformation by a converter and pulse width generation processing by a pulse width generator, obtaining the control signal of the switching tubes in the inverter module.

[0007] In some embodiments, based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor flux linkage of the motor, after being estimated by an estimator, transformed by a converter, and PI-controlled by a PI controller, the weight coefficient of the reward function in the DQN reinforcement learning process, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor are obtained, including: based on the three-phase current of the motor, being estimated by an estimator and subjected to Park transformation by a converter to obtain the sampled d-axis current of the motor, the sampled q-axis current of the motor, and the sampled rotor flux linkage of the motor; determining the absolute value of the difference between the actual speed of the motor and the target speed of the motor, denoted as the speed error of the motor; and determining the absolute value of the difference between the sampled rotor flux linkage of the motor and the target value of the rotor flux linkage of the motor, denoted as the rotor flux error of the motor; based on the speed error of the motor, after being PI-controlled by a speed-loop PI controller, obtaining the reference q-axis current of the motor; based on the rotor flux error of the motor, after being PI-controlled by a flux-loop PI controller, obtaining the reference d-axis current of the motor; and based on the speed error of the motor and the rotor flux error of the motor, after being subjected to fuzzy logic control by a fuzzy logic controller, obtaining the weight coefficient of the reward function.

[0008] In some embodiments, a fuzzy control logic is pre-set in the fuzzy logic controller; the fuzzy control logic includes: an input logic and an output logic; wherein, the input logic includes: setting the correspondence relationship between the speed error, the rotor flux error, and the error range of the input parameter; the output logic includes: setting the correspondence relationship between the weight coefficient and the output parameter range; based on the speed error of the motor and the rotor flux error of the motor, after being subjected to fuzzy logic control by a fuzzy logic controller, obtaining the weight coefficient of the reward function, including: based on the speed error of the motor and the rotor flux error of the motor, using the correspondence relationship between the set speed error, the set rotor flux error, and the error range of the input parameter, determining the error range of the input parameter corresponding to the speed error of the motor and the rotor flux error of the motor; using the error range of the input parameter corresponding to the speed error of the motor and the rotor flux error of the motor, through a fuzzy interference system and the correspondence relationship between the set weight coefficient and the output parameter range, determining the weight coefficient corresponding to the speed error of the motor and the rotor flux error of the motor as the weight coefficient of the reward function.

[0009] In some embodiments, according to the weight coefficient of the reward function, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor, after DQN reinforcement learning processing is performed by the DQN reinforcement learning algorithm, the reference d-axis voltage of the motor and the reference q-axis voltage of the motor are obtained, including: using the DQN reinforcement learning algorithm to pre-train a DQN reinforcement learning model; according to the weight coefficient of the reward function, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor, using the DQN reinforcement learning model to obtain the reference d-axis voltage of the motor and the reference q-axis voltage of the motor.

[0010] In some embodiments, using the DQN reinforcement learning algorithm to pre-train a DQN reinforcement learning model includes: collecting historical data of the motor as sample data; the historical data of the motor includes: the weight coefficient of the reward function and the historical values of the sampled d-axis current, q-axis current, reference d-axis current, reference q-axis current, reference d-axis voltage, and reference q-axis voltage of the motor; based on the DQN network, using the weight coefficient of the reward function and the historical values of the sampled d-axis current, q-axis current, reference d-axis current, and reference q-axis current of the motor as input quantities, and using the historical values of the reference d-axis voltage and reference q-axis voltage of the motor as output quantities, performing reinforcement learning training using the sample data to obtain the DQN reinforcement learning model.

[0011] In some embodiments, based on the DQN network, using the weight coefficient of the reward function and the historical values of the sampled d-axis current, q-axis current, reference d-axis current, and reference q-axis current of the motor as input quantities, and using the historical values of the reference d-axis voltage and reference q-axis voltage of the motor as output quantities, performing reinforcement learning training using the sample data to obtain the DQN reinforcement learning model includes: based on the DQN network, using the weight coefficient of the reward function and the historical values of the sampled d-axis current, q-axis current, reference d-axis current, and reference q-axis current of the motor as input quantities, and using the historical values of the reference d-axis voltage and reference q-axis voltage of the motor as output quantities, and performing reinforcement learning training using the sample data according to the following formula to obtain the DQN reinforcement learning model:

[0012] y i =r i +γmax a′ Q′(s i+1 ,a′;θ′);

[0013] where, i represents the i-th sampling amount in the sample data; y i represents the output target value, representing the historical values of the d-axis current reference quantity and the q-axis current reference quantity; r i represents the reward return value expected for reinforcement learning training using the sample data, which is a variable and uses the reward function of the DQN algorithm; γ is a constant; a′ represents the expected output action amount at the new moment, representing the historical values of the d-axis voltage reference quantity and the q-axis voltage reference quantity of the motor at the new moment; Q′ represents the initialized target network of the DQN network; θ′ represents the parameter of Q′.

[0014] Matched with the above method, on the other hand, the present invention provides a control device for a motor. The drive control end of the motor has an inverter module, and the inverter module has switching tubes. The control device for the motor includes: an acquisition unit configured to acquire the three-phase current of the motor, acquire the actual speed of the motor, acquire the target speed of the motor, and acquire the target rotor magnetic flux quantity of the motor; a control unit configured to, according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target rotor magnetic flux quantity of the motor, after performing FOC control combining fuzzy logic processing and DQN reinforcement learning processing, obtain the control signal of the switching tubes in the inverter module; the control unit is further configured to control the operation of the inverter module according to the control signal of the switching tubes in the inverter module to achieve the control of the motor.

[0015] In some embodiments, the control unit, based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor flux linkage of the motor, after performing FOC control combining fuzzy logic processing and DQN reinforcement learning processing, obtains the control signals of the switching tubes in the inverter module, including: based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor flux linkage of the motor, after being estimated by an estimator, transformed by a converter, and PI-controlled by a PI controller, obtains the weight coefficient of the reward function in the DQN reinforcement learning processing, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor; based on the weight coefficient of the reward function, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor, after performing DQN reinforcement learning processing by the DQN reinforcement learning algorithm, obtains the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor; based on the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor, after performing inverse Park transformation by a converter and pulse width generation processing by a pulse width generator, obtains the control signals of the switching tubes in the inverter module.

[0016] In some embodiments, the control unit, based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor flux linkage of the motor, after being estimated by an estimator, transformed by a converter, and PI-controlled by a PI controller, obtains the weight coefficient of the reward function in the DQN reinforcement learning processing, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor, including: based on the three-phase current of the motor, after being estimated by an estimator and Park-transformed by a converter, obtains the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, and the sampled value of the rotor flux linkage of the motor; determines the absolute value of the difference between the actual speed of the motor and the target speed of the motor, denoted as the speed error of the motor; and determines the absolute value of the difference between the sampled value of the rotor flux linkage of the motor and the target value of the rotor flux linkage of the motor, denoted as the rotor flux linkage error of the motor; based on the speed error of the motor, after performing PI control by a speed-loop PI controller, obtains the reference value of the q-axis current of the motor; based on the rotor flux linkage error of the motor, after performing PI control by a flux-linkage-loop PI controller, obtains the reference value of the d-axis current of the motor; and based on the speed error of the motor and the rotor flux linkage error of the motor, after performing fuzzy logic control by a fuzzy logic controller, obtains the weight coefficient of the reward function.

[0017] In some embodiments, the control unit is further configured to preset fuzzy control logic in a fuzzy logic controller; the fuzzy control logic includes: input logic and output logic; wherein, the input logic includes: setting the correspondence relationship between the rotational speed error, the rotor flux linkage error, and the error range of the set input parameters; the output logic includes: setting the correspondence relationship between the weight coefficient and the output parameter range; the control unit, according to the rotational speed error of the motor and the rotor flux linkage error of the motor, after performing fuzzy logic control through the fuzzy logic controller, obtains the weight coefficient of the reward function, including: according to the rotational speed error of the motor and the rotor flux linkage error of the motor, using the correspondence relationship between the set rotational speed error, the set rotor flux linkage error, and the error range of the set input parameters, determining the error range of the input parameters corresponding to the rotational speed error of the motor and the rotor flux linkage error of the motor; using the error range of the input parameters corresponding to the rotational speed error of the motor and the rotor flux linkage error of the motor, through the fuzzy interference system and the correspondence relationship between the set weight coefficient and the output parameter range, determining the weight coefficient corresponding to the rotational speed error of the motor and the rotor flux linkage error of the motor as the weight coefficient of the reward function.

[0018] In some embodiments, the control unit, according to the weight coefficient of the reward function, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor, after performing DQN reinforcement learning processing through the DQN reinforcement learning algorithm, obtains the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor, including: using the DQN reinforcement learning algorithm to pre-train a DQN reinforcement learning model; according to the weight coefficient of the reward function, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor, using the DQN reinforcement learning model to obtain the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor.

[0019] In some embodiments, the control unit pre - trains a DQN reinforcement learning model using the DQN reinforcement learning algorithm, including: collecting historical data of the motor as sample data; the historical data of the motor includes: the weight coefficient of the reward function, and historical values of the sampled d - axis current, sampled q - axis current, d - axis current reference, q - axis current reference, d - axis voltage reference, and q - axis voltage reference of the motor; based on the DQN network, using the weight coefficient of the reward function and historical values of the sampled d - axis current, sampled q - axis current, d - axis current reference, and q - axis current reference of the motor as input quantities, and using historical values of the d - axis voltage reference and q - axis voltage reference of the motor as output quantities, performing reinforcement learning training using the sample data to obtain the DQN reinforcement learning model.

[0020] In some embodiments, the control unit, based on the DQN network, uses the weight coefficient of the reward function and historical values of the sampled d - axis current, sampled q - axis current, d - axis current reference, and q - axis current reference of the motor as input quantities, and uses historical values of the d - axis voltage reference and q - axis voltage reference of the motor as output quantities, and performs reinforcement learning training using the sample data to obtain the DQN reinforcement learning model, including: based on the DQN network, using the weight coefficient of the reward function and historical values of the sampled d - axis current, sampled q - axis current, d - axis current reference, and q - axis current reference of the motor as input quantities, and using historical values of the d - axis voltage reference and q - axis voltage reference of the motor as output quantities, and performing reinforcement learning training using the sample data according to the following formula to obtain the DQN reinforcement learning model:

[0021] y i =r i +γmax a′ Q′(s i+1 ,a′;θ′);

[0022] Where i represents the i - th group of sampled quantities in the sample data; y i represents the output target value, representing historical values of the d - axis current reference and q - axis current reference; r i represents the expected reward return value for performing reinforcement learning training using the sample data, which is a variable and uses the reward function of the DQN algorithm; γ is a constant; a′ represents the expected output action quantity at the new moment, representing historical values of the d - axis voltage reference and q - axis voltage reference of the motor at the new moment; Q′ represents the initialized target network of the DQN network; θ′ represents the parameter of Q′.

[0023] Matched with the above - mentioned device, on the other hand, the present invention provides a motor, including: the control device of the motor described above.

[0024] Matched with the above device, on the other hand, the present invention provides a compressor, comprising: the control device of the motor described above, or the motor described above.

[0025] Matched with the above device, on the other hand, the present invention provides an air conditioner, comprising: the control device of the motor described above, or the motor described above, or the compressor described above.

[0026] Matched with the above method, on the other hand, the present invention provides a storage medium, the storage medium comprising a stored program, wherein, when the program runs, it controls the device where the storage medium is located to execute the steps of the control method of the motor described above.

[0027] Matched with the above method, on the other hand, the present invention provides a computer program product, comprising a computer program, which when executed by a processor implements the steps of the control method of the motor described above.

[0028] Thus, the solution of the present invention, for the motor drive control mode (such as the motor drive control mode in an air conditioner compressor) and the inverter used thereby (such as a two-level three-phase inverter), samples the three-phase current of the motor (such as the three-phase current I a 、I b and I c ) output by the two-level three-phase inverter), samples the actual speed of the motor (such as the feedback sampling quantity ω fb ) of the rotor angular velocity of the motor, and obtains the target speed of the motor (such as the given rotor angular velocity reference quantity ω ref ) of the motor, as well as the target quantity of the rotor magnetic flux of the motor (such as the rotor magnetic flux reference quantity |ψ r_ref |); according to the three-phase current of the motor (such as the three-phase current I a 、I b and I c ) output by the two-level three-phase inverter, the d-axis current sampling quantity I d and the q-axis current sampling quantity I q 、the rotor magnetic flux sampling quantity (the rotor magnetic flux sampling quantity |ψ r |) of the motor are obtained through the processing of the estimator and the converter; the absolute value of the difference in rotor speed (i.e., the rotor angular velocity error |ω r_error |) between the actual speed of the motor (such as the feedback sampling quantity ω fb ) and the target speed of the motor (such as the given rotor angular velocity reference quantity ω ref ) is obtained through the speed loop PI controller to obtain the q-axis current reference quantity I qref , the rotor magnetic flux sampling quantity (the rotor magnetic flux sampling quantity |ψ r |) of the motor, the target quantity of the rotor magnetic flux of the motor (such as the rotor magnetic flux reference quantity |ψ r_ref|), the absolute value of the rotor flux linkage difference (i.e., the rotor flux error |ψ r_error |) passes through the flux linkage loop PI controller to obtain the d-axis current reference I dref , the absolute value of the rotor speed difference (i.e., the rotor angular velocity error |ω r_error |), the absolute value of the rotor flux linkage difference (i.e., the rotor flux error |ψ r_error |) passes through the fuzzy logic controller to obtain the weight coefficient λ; the weight coefficient λ, the q-axis current reference I qref , the q-axis current sampled value I q , the d-axis current reference I dref , the d-axis current sampled value I d , after being processed by DQN reinforcement learning, the d-axis voltage reference V dref and the q-axis voltage reference V qref are obtained; the d-axis voltage reference V dref and the q-axis voltage reference V qref , after passing through the converter and pulse width modulation processing, obtain the control signals of each switch tube in the inverter to control the inverter to realize the drive control of the motor; thus, in the motor drive control, based on the three-phase current and speed of the motor, adopting the FOC control combining fuzzy logic processing and DQN reinforcement learning processing to control the inverter to realize the drive control of the motor and improve the control effect of the motor.

[0029] Other features and advantages of the present invention will be described in the subsequent description, and part of them will be obvious from the description or understood by implementing the present invention.

[0030] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0031] Figure 1 is a schematic flowchart of an embodiment of the control method of the motor of the present invention;

[0032] Figure 2 is a schematic flowchart of an embodiment of obtaining the control signals of the switch tubes in the inverter module after being estimated by the estimator, transformed by the converter, PI controlled, fuzzy logic controlled, DQN reinforcement learning processed and pulse width modulated in the method of the present invention;

[0033] Figure 3 is a schematic flowchart of an embodiment of obtaining the weight coefficient of the reward function, the d-axis current sampled value, the q-axis current sampled value, the d-axis current reference and the q-axis current reference after being estimated by the estimator, transformed by the converter and PI controlled by the PI controller in the method of the present invention;

[0034] Figure 4Schematic flowchart of an embodiment for obtaining the weight coefficient of the reward function after fuzzy logic control by the fuzzy logic controller in the method of the present invention;

[0035] Figure 5 Schematic flowchart of an embodiment for obtaining the d-axis voltage reference and q-axis voltage reference of the motor after DQN reinforcement learning processing by the DQN reinforcement learning algorithm in the method of the present invention;

[0036] Figure 6 Schematic flowchart of an embodiment for pre-training a DQN reinforcement learning model using the DQN reinforcement learning algorithm in the method of the present invention;

[0037] Figure 7 Schematic structural diagram of an embodiment of the control device for the motor of the present invention;

[0038] Figure 8 Schematic block diagram of the principle of the DQN algorithm;

[0039] Figure 9 Schematic flowchart of the loop training process of the DQN algorithm;

[0040] Figure 10 Schematic block diagram of the overall system of FOC control PMSM based on DQN reinforcement learning optimized by fuzzy logic;

[0041] Figure 11 Schematic structural diagram of the main circuit of a two-level three-phase inverter circuit;

[0042] Figure 12 Curve schematic diagram of the input function of the fuzzy logic controller (such as the relationship between the rotor flux linkage error |ψ r_error | and VS, S, M, L, VL);

[0043] Figure 13 Curve schematic diagram of the input function of the fuzzy logic controller (such as the relationship between the rotor angular velocity error |ω r_error | and VS, S, M, L, VL);

[0044] Figure 14 Curve schematic diagram of the output function of the fuzzy logic controller (such as the relationship between λ and NL, NM, NS, O, PS, PM, PL);

[0045] Figure 15 Schematic flowchart of the motor drive control method in an air conditioner compressor based on DQN reinforcement learning magnetic field orientation control optimized by fuzzy logic.

[0046] In combination with the accompanying drawings, the reference numerals in the embodiments of the present invention are as follows:

[0047] 102 - Acquisition unit; 104 - Control unit. Detailed implementation manner

[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0049] Considering that in motor drive control (such as motor drive control in an air-conditioning compressor), when using FOC to control a three-phase inverter in related solutions, two PI controllers are required for the d-axis control loop and the q-axis control loop, and there is a problem of difficult parameter tuning under different power supplies and loads, which affects the control effect of the motor.

[0050] The Deep Q-Network (DQN) algorithm is a method that combines deep learning and reinforcement learning. It solves the decision-making problem in a high-dimensional state space by approximating the Q-value function through a neural network. It adopts an experience replay mechanism and a target network, which significantly improves the training stability and efficiency. Applying DQN to a three-phase inverter is a very promising research direction, and DQN can optimize the performance of the three-phase inverter through intelligent decision-making.

[0051] In addition, reinforcement learning (RL) is a machine learning method. As a popular research field of artificial intelligence technology, reinforcement learning models and explores the external environment, and can obtain an optimized reward learning strategy through continuous adjustment and update. In this process, only limited information is required to complete. Applying it to a three-phase inverter can significantly improve the energy efficiency, stability and performance of the system through its ability to handle complex operations and complex control problems.

[0052] Therefore, the solution of the present invention proposes a control method for a motor, specifically a motor drive control method for an air-conditioning compressor based on DQN reinforcement learning magnetic field orientation control optimized by fuzzy logic. It uses the DQN reinforcement learning algorithm to replace the PI controllers of the d-axis and q-axis in the FOC algorithm, and retains the PI controllers of the speed loop and the magnetic flux loop, which can avoid the problem of difficult parameter tuning under different power supplies and loads and improve the control effect of the motor.

[0053] According to an embodiment of the present invention, a control method for a motor is provided, such as Figure 1Schematic flow diagram of an embodiment of the method of the present invention. The drive control terminal of the motor has an inverter module (such as a two-level three-phase inverter), and the inverter module has switching tubes; the inverter module is used to perform an inversion process on the input DC voltage source under the control of the control signals of the switching tubes to drive and control the motor. Figure 11 Is the structural schematic diagram of the main circuit of a two-level three-phase inverter circuit. As Figure 11 Shown, the main circuit of the two-level three-phase inverter circuit includes: DC bus voltage V dc , a DC bus capacitor formed by connecting capacitor C1 and capacitor C2 in series, and switching tubes S a , switching tube S b , switching tube S c , switching tube Switching tube Switching tube To form a three-phase inverter, and a permanent magnet synchronous motor (PMSM).

[0054] Among them, the positive pole of the DC bus voltage V dc Is connected to the positive pole of capacitor C1, and the negative pole of capacitor C1 is connected to the positive pole of capacitor C2; the negative pole of the DC bus voltage V dc Is connected to the negative pole of capacitor C2, and the midpoint of capacitor C1 and capacitor C2 is point o. The positive pole of the DC bus voltage V dc Is also respectively connected to the first connection ends of each switching tube in switching tube S a , switching tube S b , switching tube S c ; the negative pole of the DC bus voltage V dc Is also respectively connected to the second connection ends of each switching tube in switching tube Switching tube Switching tube ; the second connection ends of each switching tube in switching tube S a , switching tube S b , switching tube S c Are correspondingly connected to the first connection ends of each switching tube in switching tube Switching tube Switching tube . The common end of the second connection end of switching tube S a And the first connection end of switching tube Is point a, and point a is connected to the first connection end of the first-phase winding in the three-phase windings of the permanent magnet synchronous motor (PMSM); the common end of the second connection end of switching tube S b And the first connection end of switching tube Is point b, and point b is connected to the first connection end of the second-phase winding in the three-phase windings of the permanent magnet synchronous motor (PMSM); the common end of the second connection end of switching tube S c And the first connection end of switching tube The common terminal of the first connection end is point c, and point c is connected to the first connection end of the third-phase winding in the three-phase winding of the permanent magnet synchronous motor (PMSM). The second connection ends of the first-phase winding, the second-phase winding, and the third connection end of the third-phase winding are all connected to point n.

[0055] In the solution of the present invention, as Figure 1 shown, the control method of the motor includes: step S110 to step S130.

[0056] At step S110, when the motor is running, obtain the three-phase current of the motor, obtain the actual speed of the motor, obtain the target speed of the motor, and obtain the target value of the rotor magnetic flux of the motor; wherein, the three-phase current of the motor is like the three-phase current I a 、I b and I c output by the two-level three-phase inverter, the actual speed of the motor is like the feedback sampling quantity ω fb of the rotor angular velocity of the motor, the target speed of the motor is like the given reference quantity ω ref of the rotor angular velocity of the motor, and the target value of the rotor magnetic flux of the motor is like the reference quantity |ψ r_ref | of the rotor magnetic flux. Among them, since for the star-connected winding, the sum of the three-phase currents of the motor is 0, that is, Ia + Ib + Ic = 0. Therefore, for the star-connected winding, obtaining the three-phase current of the motor is also obtaining the two-phase current of the motor, and the third-phase current of the motor can be calculated from the two-phase current of the motor.

[0057] At step S120, according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, after performing FOC control combining fuzzy logic processing and DQN reinforcement learning processing, obtain the control signal of the switching tube in the inverter module; specifically, according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, after being estimated by an estimator, transformed by a converter, PI controlled, fuzzy logic controlled, DQN reinforcement learning processed, and pulse width modulated, obtain the control signal of the switching tube in the inverter module. Among them, the FOC control combining fuzzy logic processing and DQN reinforcement learning processing includes: estimator estimation, converter transformation, PI control, fuzzy logic control, DQN reinforcement learning processing, and pulse width modulation processing. The converter is used to perform Park transformation and inverse Park transformation. PI control is the control using a PI controller, fuzzy logic control is the control using a fuzzy logic controller, DQN reinforcement learning processing is the processing using a DQN reinforcement learning algorithm, and pulse width modulation processing is the processing using a pulse width generator.

[0058] At step S130, according to the control signals of the switching tubes in the inverter module, control the inverter module to operate to achieve the control of the motor, specifically to achieve the drive control of the motor.

[0059] For the DQN reinforcement learning algorithm proposed in the solution of the present invention, the reinforcement learning algorithm has the ability to handle complex operation and complex control problems, which can significantly improve the energy efficiency, stability and performance of the system; the DQN reinforcement learning algorithm interacts with the environment and adjusts the decisive parameters based on the expected reward, and finally achieves the role of target control by outputting the most appropriate action parameters. DQN can dynamically adjust the working mode of the three-phase inverter according to the real-time load conditions and power supply characteristics to achieve a higher energy efficiency ratio.

[0060] In some embodiments, the specific process of obtaining the control signals of the switching tubes in the inverter module after adopting FOC control combining fuzzy logic processing and DQN reinforcement learning processing according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor in step S120 is as follows in the following exemplary description.

[0061] The following combines Figure 2 A schematic flowchart of an embodiment of the control signals of the switching tubes in the inverter module obtained after being estimated by an estimator, transformed by a converter, PI controlled, fuzzy logic controlled, DQN reinforcement learning processed, and pulse width modulated in the method of the present invention shown below further illustrates the specific process of obtaining the control signals of the switching tubes in the inverter module after being estimated by an estimator, transformed by a converter, PI controlled, fuzzy logic controlled, DQN reinforcement learning processed, and pulse width modulated in step S120, including: step S210 to step S230.

[0062] Step S210, according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, after being estimated by an estimator, transformed by a converter, and PI controlled by a PI controller, obtain the weight coefficient of the reward function in the DQN reinforcement learning process, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor; wherein, the weight coefficient of the reward function is like the weight coefficient λ, the sampled value of the d-axis current of the motor is like the sampled value of the d-axis current I d , the sampled value of the q-axis current of the motor is like the sampled value of the q-axis current I q , the reference value of the d-axis current of the motor is like the reference value of the d-axis current I dref , and the reference value of the q-axis current of the motor is like the reference value of the q-axis current I qref .

[0063] Step S220: Based on the weight coefficient of the reward function, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor, after performing DQN reinforcement learning processing through the DQN reinforcement learning algorithm, the reference d-axis voltage of the motor and the reference q-axis voltage of the motor are obtained. The reference d-axis voltage of the motor is like the reference d-axis voltage V dref , and the reference q-axis voltage of the motor is like the reference q-axis voltage V qref .

[0064] Step S230: Based on the reference d-axis voltage of the motor and the reference q-axis voltage of the motor, after performing inverse Park transformation through a converter and pulse width generation processing through a pulse width generator, the control signals of the switching tubes in the inverter module are obtained.

[0065] A motor drive control scheme for an air conditioner compressor based on DQN reinforcement learning magnetic field orientation control optimized by fuzzy logic proposed by the solution of the present invention uses the DQN reinforcement learning algorithm to replace the PI controllers of the d-axis and q-axis in the FOC algorithm, retains the PI controllers of the speed loop and the magnetic flux loop, and then according to the rotor magnetic flux error amount |ψ r_error | and the rotor angular velocity error amount |ω r_error |, uses fuzzy logic judgment to quickly lock the corresponding weight coefficient, and applies the weight coefficient λ to the previous reward function r t () to obtain a new reward function, realizing the decoupled control of the rotor magnetic flux ψ r and the rotational speed (such as the rotor angular velocity ω fb ), making the calculation more convenient and the control more efficient.

[0066] In some embodiments, in step S210, the specific process of obtaining the weight coefficient of the reward function, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor in the DQN reinforcement learning processing through estimation by an estimator, transformation by a converter, and PI control by a PI controller based on the three-phase current of the motor, the actual rotational speed of the motor, the target rotational speed of the motor, and the target rotor magnetic flux of the motor is as follows in the following exemplary description.

[0067] The following combines with Figure 3Schematic diagram of an embodiment process of obtaining the weight coefficient of the reward function, the sampled d-axis current, the sampled q-axis current, the reference d-axis current, and the reference q-axis current after estimation by an estimator, transformation by a converter, and PI control by a PI controller in the method of the present invention, further illustrating the specific process of obtaining the weight coefficient of the reward function, the sampled d-axis current, the sampled q-axis current, the reference d-axis current, and the reference q-axis current after estimation by an estimator, transformation by a converter, and PI control by a PI controller in step S210, including: steps S310 to S330.

[0068] Step S310, based on the three-phase current of the motor, estimate through an estimator and perform Park transformation through a converter to obtain the sampled d-axis current of the motor, the sampled q-axis current of the motor, and the sampled rotor magnetic flux of the motor; wherein, the sampled d-axis current of the motor is like the sampled d-axis current I d , the sampled q-axis current of the motor is like the sampled q-axis current I q , and the sampled rotor magnetic flux of the motor is like the sampled rotor magnetic flux |ψ r |.

[0069] Step S320, determine the absolute value of the difference between the actual speed of the motor and the target speed of the motor, denoted as the speed error of the motor, such as the rotor angular velocity error |ω r_error |; and determine the absolute value of the difference between the sampled rotor magnetic flux of the motor and the target rotor magnetic flux of the motor, denoted as the rotor magnetic flux error of the motor, such as the rotor magnetic flux error |ψ r_error |.

[0070] Step S330, based on the speed error of the motor, after performing PI control through a speed-loop PI controller, obtain the reference q-axis current of the motor, such as the reference q-axis current I qref ; based on the rotor magnetic flux error of the motor, after performing PI control through a magnetic-flux-loop PI controller, obtain the reference d-axis current of the motor, such as the reference d-axis current I dref ; and based on the speed error of the motor and the rotor magnetic flux error of the motor, after performing fuzzy logic control through a fuzzy logic controller, obtain the weight coefficient of the reward function, specifically obtain the weight coefficient of the reward function r t in DQN reinforcement learning, such as the weight coefficient λ.

[0071] In the solution of the present invention, based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target rotor magnetic flux of the motor, retain the PI controllers of the speed loop and the magnetic-flux loop in FOC control, and further based on the rotor magnetic flux error amount |ψ r_error | and the rotor angular velocity error amount |ωr_error |, the fuzzy logic judgment is used to quickly lock the corresponding weight coefficient, and the weight coefficient λ is applied to the previous reward function r t () to obtain a new reward function, and realize the decoupling control of the rotor magnetic flux ψ r and the rotational speed (such as the rotor angular velocity ω fb ), making the calculation more convenient and the control more efficient.

[0072] In some embodiments, a fuzzy control logic is preset in the fuzzy logic controller; the fuzzy control logic includes: an input logic and an output logic; wherein, the input logic includes: setting the corresponding relationship between the rotational speed error, the rotor magnetic flux error, and the set input parameter error range, such as Figure 12 、 Figure 13 and the relationship reflected in Table 1; the output logic includes: setting the corresponding relationship between the weight coefficient and the set output parameter range, such as Figure 14 reflected. Among them, 1e-4 represents 10 -4 , 3e-4 represents 30 -4 , 5e-4 represents 50 -4 , 7e-4 represents 70 -4 , 1e-3 represents 10 -3 .

[0073] The specific process of obtaining the weight coefficient of the reward function through fuzzy logic control by the fuzzy logic controller according to the rotational speed error of the motor and the rotor magnetic flux error of the motor in step S330 is shown in the following exemplary description.

[0074] The following combines Figure 4 a schematic flowchart of an embodiment of the weight coefficient of the reward function obtained through fuzzy logic control by the fuzzy logic controller in the method of the present invention shown, and further illustrates the specific process of obtaining the weight coefficient of the reward function through fuzzy logic control by the fuzzy logic controller in step S330, including: step S410 to step S420.

[0075] Step S410, according to the rotational speed error of the motor and the rotor magnetic flux error of the motor, using the corresponding relationship between the set rotational speed error, the set rotor magnetic flux error, and the set input parameter error range, determine the input parameter error range corresponding to the rotational speed error of the motor and the rotor magnetic flux error of the motor.

[0076] Step S420: Determine the weight coefficient corresponding to the rotational speed error and the rotor flux error of the motor by using the fuzzy interference system and the corresponding relationship between the set weight coefficient and the set output parameter range, and use it as the weight coefficient of the reward function.

[0077] Figure 12 It is a schematic curve diagram of the input function of the fuzzy logic controller (such as the relationship between the rotor flux error |ψ r_error | and VS, S, M, L, VL). Figure 13 It is a schematic curve diagram of the input function of the fuzzy logic controller (such as the relationship between the rotor angular velocity error |ω r_error | and VS, S, M, L, VL). The input function of the fuzzy logic controller (such as the relationship between the rotor flux error |ψ r_error | and VS, S, M, L, VL), the input function of the fuzzy logic controller (such as the relationship between the rotor angular velocity error |ω r_error | and VS, S, M, L, VL), as Figure 12 and Figure 13 shown, ultimately aims to output VS, S, M, L, VL. In the cross-region of the above-mentioned quantities caused by different values, generally take the larger value among VS, S, M, L, VL. In addition, the input of the fuzzy logic should select the most suitable input parameter error range according to the rotor flux error quantity such as the rotor angular velocity error |ω r_error | and the rotor angular velocity error quantity such as the rotor flux error |ψ r_error |. Here, the rotor flux error range is set as [0, 0.001], with the unit of wb (Weber), as Figure 12 shown; the rotor angular velocity error range is [0, 3.5], with the unit of rad / s (radian per second), as Figure 13 shown. That is, when each error falls within the above range, the corresponding weight coefficient λ can be quickly determined through the fuzzy interference system (Fuzzy Inference System, FIS). Here, it should be noted that when the rotor flux error is less than 0 or the current error is greater than 0.001, its input amplitude needs to be limited within the range of [0, 0.001], and the same applies to the rotor angular velocity error range. That is:

[0078]

[0079] Figure 14Schematic diagram of the curve of the output function of the fuzzy logic controller (such as the relationship between λ and NL, NM, NS, O, PS, PM, PL). Among them, for the selection of each quantity of NL, NM, NS, O, PS, PM, PL, refer to Table 1. After selecting each quantity of NL, NM, NS, O, PS, PM, PL, then according to Figure 14 Find the approximate λ in a similar way to the selection of VS, S, M, L, VL. Specifically, when each quantity in the above equation is edited into the fuzzy interference system, it will output specific weight coefficients without listing the specific expression function. In the control method proposed by the solution of the present invention, a fuzzy logic controller is used to achieve the decoupling of the rotor magnetic flux ψ r and the rotor angular velocity ω fb . The error quantities of the above two parameters, namely the rotor angular velocity error |ω r_error | and the rotor magnetic flux error |ψ r_error |, are used as the inputs of the fuzzy logic controller. After passing through the FIS, a weight coefficient λ is output. The range of this coefficient is 0 < λ ≤ 1. As shown in Figure 14 , the principle of selecting the weight coefficient λ is shown in Table 1.

[0080] Table 1: Action rules of the fuzzy logic controller

[0081]

[0082] Figure 14 The parameters in represent the value range of the weight coefficient λ, which is obtained by taking fuzzy judgment based on the magnetic flux error and angular velocity error collected in Table 1. The parameter that the fuzzy logic controller needs to obtain is the weight coefficient λ in Figure 14 .

[0083] Table 1 and Figure 12 , Figure 13 The horizontal axis in is the absolute value of the error quantity (hereinafter simply referred to as the error). VS, S, M, L, VL represent the fuzzy judgment values when the error is located in different ranges. VS = Very Small (very small), S = Small (small), M = medium (medium), L = Large (large), VL = Very Large (very large). The character meanings in the table: L is the abbreviation of Large (large), M is the abbreviation of Medium (medium), S is the abbreviation of Small (small), P is the abbreviation of Positive (positive), and N is the abbreviation of Negative (negative).

[0084] In related solutions, when FOC control is adopted, quantities output by the motor, such as rotor magnetic flux and rotor speed, are collected. Due to a certain degree of coupling between the above parameters, there is an interference problem during control. Therefore, the solution of the present invention proposes a fuzzy logic controller, which can complete the decoupling control of motor parameters by fuzzifying the rotor magnetic flux and speed. Aiming at the problem of strong coupling of motor parameters, the solution of the present invention uses a fuzzy logic controller to achieve a certain degree of decoupling between the motor speed (such as rotor angular velocity ω fb ) and rotor magnetic flux ψ r , realizes the decoupling of rotor magnetic flux error and rotor angular velocity error, and achieves the purpose of better controlling the motor.

[0085] In some embodiments, in step S220, according to the weight coefficient of the reward function, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor, after performing DQN reinforcement learning processing through the DQN reinforcement learning algorithm, the specific process of obtaining the reference d-axis voltage of the motor and the reference q-axis voltage of the motor is as follows in the following exemplary description.

[0086] The following combines Figure 5 A schematic flowchart of an embodiment of the method of the present invention for obtaining the reference d-axis voltage of the motor and the reference q-axis voltage of the motor after performing DQN reinforcement learning processing through the DQN reinforcement learning algorithm further illustrates the specific process of obtaining the reference d-axis voltage of the motor and the reference q-axis voltage of the motor after performing DQN reinforcement learning processing in step S220, including: step S510 to step S520.

[0087] Step S510, using the DQN reinforcement learning algorithm, pre-train a DQN reinforcement learning model, such as the agent module of the DQN algorithm.

[0088] Step S520, according to the weight coefficient of the reward function, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor, use the DQN reinforcement learning model to obtain the reference d-axis voltage of the motor and the reference q-axis voltage of the motor.

[0089] Figure 10 It is a schematic block diagram of the overall system of FOC control PMSM based on DQN reinforcement learning optimized by fuzzy logic. In the FOC algorithm in related solutions, the d-axis control loop and the q-axis control loop separately use two controllers (i.e., two PI controllers); in Figure 10In the example shown, the DQN reinforcement learning algorithm is used to replace the PI controllers on the d-axis and q-axis in the FOC algorithm, while the PI controllers of the speed loop and the flux linkage loop are retained. In Figure 10 the example shown, it includes the PI controller of the speed loop and the PI controller of the flux linkage loop.

[0090] Figure 10 where, |ω r_error | represents the rotor angular velocity error; |ψ r_error | represents the rotor flux linkage error; ω ref represents the rotor angular velocity reference; I qref represents the q-axis current reference; I q represents the q-axis current sampling value; |ψ r | represents the rotor flux linkage sampling value; |ψ r_ref | represents the rotor flux linkage reference; I dref represents the d-axis current reference; I d represents the d-axis current sampling value; V qref represents the q-axis voltage reference; V dref represents the d-axis voltage reference; V αref represents the α-axis voltage reference; V βref represents the β-axis voltage reference; S a , S b , S c represents the switching control quantities of phases a, b, and c; V dc represents the voltage of the DC voltage source; I a represents the a-phase current sampling value; I b represents the b-phase current sampling value; I α represents the α-axis current sampling value; I β represents the β-axis current sampling value; θ e represents the converter a-phase angle measurement; ω fb represents the rotor angular velocity feedback sampling value.

[0091] The overall block diagram of the FOC control of PMSM with fuzzy logic DQN is as shown in Figure 10 According to Figure 10 it can be seen that among the three-phase currents I a , I b and I c output by the two-level three-phase inverter, the a-phase current sampling value I a and the b-phase current sampling value I b , through the estimator, the α-axis current sampling value I α , the β-axis current sampling value I β , the converter a-phase angle measurement θ e , and the rotor flux linkage sampling value |ψ r | are obtained. Based on the converter a-phase angle measurement θ e, the α-axis current sampling value I α , the β-axis current sampling value I β After Park transformation, the output d-axis current sampling value I d and the q-axis current sampling value I q . Estimators are core tools in statistics and machine learning for parameter estimation or model construction, which infer population characteristics or model parameters through sample data. The d-axis current reference value I dref is obtained from the rotor flux sampling value |ψ r | and the rotor flux reference value |ψ r_ref |. The rotor flux error |ψ r_error | obtained after comparison is processed by a PI controller. The q-axis current reference value I qref is obtained from the rotor angular velocity feedback sampling value ω fb and the given rotor angular velocity reference value ω ref . The rotor angular velocity error |ω r_error | obtained after comparison is processed by a PI controller. Park transformation (i.e., Park's transformation) projects physical quantities such as current and voltage in the three-phase stationary coordinate system (abc) onto the rotating coordinate system (dq0) composed of the direct axis (d-axis), quadrature axis (q-axis), and zero axis (0-axis) that rotates with the rotor, thereby simplifying the dynamic model of the motor. The inverse Park transformation is a mathematical method for converting variables in the two-phase rotating coordinate system (dq coordinate system) back to the two-phase stationary coordinate system (αβ coordinate system), mainly used in motor control and power system analysis.

[0092] Among them, the rotor angular velocity error |ω r_error | and the rotor flux error |ψ r_error | are respectively used as the inputs of the fuzzy logic controller. The fuzzy logic controller outputs a weight coefficient λ to the DQN agent (i.e., the DQN reinforcement learning agent), and then the DQN agent outputs the d-axis voltage reference value V dref and the q-axis voltage reference value V qref . After the inverse Park transformation of this reference value (i.e., the d-axis voltage reference value V dref and the q-axis voltage reference value V qref ), the control signals S a , S b and S c for controlling the switching tubes in the two-level three-phase inverter can be generated through the pulse width generation module. The two-level three-phase inverter outputs three-phase voltages V a , V b and V c to the permanent magnet synchronous motor PMSM.

[0093] In related solutions, when using FOC to control a three-phase inverter, two PI controllers are required for the d-axis control loop and the q-axis control loop. There are problems with difficult parameter tuning for the two PI controllers under various power supplies and loads. However, in the solution of the present invention, after applying reinforcement learning technology, it is not necessary to complete the parameter tuning of the above control loops, and its self-learning technology can be used to complete the adaptation and reference parameter output in specific situations. Aiming at the problem of difficult parameter tuning for the PI of the d-axis and q-axis controllers, the solution of the present invention uses the DQN reinforcement learning method to automatically tune the corresponding weight coefficients and bias amounts according to the parameters using the backpropagation algorithm, which is dynamically better and has the capabilities of self-learning and self-correction. Among them, the reinforcement learning includes a backpropagation neural network, and there is a bias amount in each neuron of the backpropagation neural network. The weight coefficients described here are not the weight coefficients λ of the fuzzy logic control in the text. In the backpropagation neural network, each neuron has a corresponding weight coefficient and bias amount. Generally speaking, the weight coefficients in the neural network are represented by w (weight), and the bias amount is represented by b (bias).

[0094] In some embodiments, the specific process of pre-training the DQN reinforcement learning model using the DQN reinforcement learning algorithm in step S510 is as follows in the following exemplary description.

[0095] The following combines Figure 6 FIG. shows a schematic flowchart of an embodiment of pre-training the DQN reinforcement learning model using the DQN reinforcement learning algorithm in the method of the present invention, and further illustrates the specific process of pre-training the DQN reinforcement learning model using the DQN reinforcement learning algorithm in step S510, including: steps S610 to S620.

[0096] Step S610, collecting historical data of the motor as sample data; the historical data of the motor includes: the weight coefficients of the reward function, and the historical values of the d-axis current sampling amount, q-axis current sampling amount, d-axis current reference amount, q-axis current reference amount, d-axis voltage reference amount, and q-axis voltage reference amount of the motor.

[0097] Step S620, based on the DQN network, using the weight coefficients of the reward function, and the historical values of the d-axis current sampling amount, q-axis current sampling amount, d-axis current reference amount, and q-axis current reference amount of the motor as input quantities, and using the historical values of the d-axis voltage reference amount and q-axis voltage reference amount of the motor as output quantities, performing reinforcement learning training using the sample data to obtain the DQN reinforcement learning model.

[0098] Figure 8 is a schematic block diagram of the principle of the DQN algorithm. In Figure 8 The experience replay pool shown contains the content of two steps, one is storage and the other is replay. Storage means using the experience tuple (st , a t , r t+1 , s t+1 ) and other data are stored as experience in the experience replay pool; replay means collecting experience data from the experience replay pool through certain specific rules during the training process.

[0099] Figure 9 It is a schematic flow diagram of the loop training process of the DQN algorithm. Figure 8 and Figure 9 The environment in, which represents the scenario where the DQN algorithm is applied. For example, the application scenario in the solution of the present invention is the motor drive control scenario. Among them, the environment here refers to the environment other than DQN reinforcement learning, mainly determined by specific sampling quantities such as (λ, I qref , I q , I dref , I d ) and other environmental parameters. See Figure 8 , Figure 9 and Figure 10 In the example shown, in the DQN algorithm, the state s includes five quantities input to the DQN reinforcement learning agent, namely s0 represents the state at the initial moment, and generally the above values are set to 0. a is the output action. According to Figure 10 , it can be known that a = (V qref , V dref ). I qref represents the q-axis current reference quantity, I q represents the q-axis current sampling quantity, I dref represents the d-axis current reference quantity, I d represents the d-axis current sampling quantity, and λ represents a constant weight coefficient.

[0100] In the solution of the present invention, the DQN reinforcement learning algorithm is used to achieve the role of target control by outputting the most suitable action parameters, so as to dynamically adjust the working mode of the three-phase inverter according to the real-time load conditions and power supply characteristics, in order to achieve a higher energy efficiency ratio.

[0101] In some embodiments, in step S620, based on the DQN network, using the weight coefficient of the reward function, and the historical values of the sampled d-axis current, sampled q-axis current, reference d-axis current, and reference q-axis current of the motor as input quantities, and the historical values of the reference d-axis voltage and reference q-axis voltage of the motor as output quantities, the sample data is used for reinforcement learning training to obtain the DQN reinforcement learning model, including: based on the DQN network, using the weight coefficient of the reward function, and the historical values of the sampled d-axis current, sampled q-axis current, reference d-axis current, and reference q-axis current of the motor as input quantities, and the historical values of the reference d-axis voltage and reference q-axis voltage of the motor as output quantities, according to the following formula, the sample data is used for reinforcement learning training to obtain the DQN reinforcement learning model:

[0102] y i =r i +γmax a′ Q′(s i+1 ,a′;θ′);

[0103]

[0104] where i represents the i-th group of sampled quantities in the sample data; y i represents the output target value, representing the historical values of the reference d-axis current and reference q-axis current; r i represents the reward return value expected from the reinforcement learning training using the sample data, which is a variable, using the reward function of the DQN algorithm, and the reward function of the DQN algorithm is like the function r t (), t represents the time step; γ is a constant; a represents the expected output action quantity, representing the historical values of the reference d-axis voltage and reference q-axis voltage of the motor, a′ represents the expected output action quantity at the new moment, representing the historical values of the reference d-axis voltage and reference q-axis voltage of the motor at the new moment; Q represents the initialized network of the DQN network, Q′ represents the initialized target network of the DQN network; θ′ represents the parameters of Q′; Q1 represents the Q-value function of the direct-axis current error i derror , Q2 represents the Q-value function of the quadrature-axis current error i qerror ; μ j is the action function parameter of the previous time step, j represents the number of actions, and τ is a constant.

[0105] As Figure 8 and Figure 9 shown, the process flow of the cyclic training process of the DQN algorithm includes:

[0106] Step 11. Initialization: Initialize the experience replay buffer D; Initialize the DQN network Q(s, a; θ); Initialize the target network Q′(s, a; θ′); Initialize the state s0.

[0107] Among them, for initialization, it only needs to assign an initial value to each variable before the start of the program. Initialization is a crucial programming step, and its purpose is to ensure that variables, objects, and systems are in the expected initial state before the program runs. The experience replay buffer D is a buffer in the experience replay pool.

[0108] Among them, t represents the time step. s represents the state, s t represents the state at the current time step t, s t+1 represents the state at the next time step t + 1; a represents the action, a t represents the action at the current time step t, a t+1 represents the action at the next time step t + 1; θ represents the parameters involved in the update and calculation process, such as the parameters of the initialized DQN network Q; θ′ represents the parameters of the initialized target network Q′ (i.e., Target Q). r represents the reward, r t represents the reward at the current t, r t+1 represents the reward at the next time step t + 1.

[0109] Step 12. Loop training:

[0110] For each time step t: The time step t is each calculation cycle of the DQN network. Generally speaking, the smaller t is set, the more accurate the output control amount is, and the corresponding calculation complexity is greater. Determine the time step t to complete the loop update of the DQN network.

[0111] Step 121. Select an action: Select the current action a t according to the current state s t and the current policy (such as the ε-greedy policy);

[0112] Step 122. Execute the action: Execute the current action a t , observe the next state s t+1 and the reward r t .

[0113] Step 123. Store the experience: Store the experience tuple (s t , a t , r t , s t+1 ) into the experience replay buffer D.

[0114] Step 124. Sample the experience: Randomly sample a batch of experience tuples (s i , a i , ri , s i+1 ). i represents the i-th group of sampling amounts.

[0115] Step 125, calculate the target value:

[0116] y i = r i + γmax a′ Q′(s i+1 , a′; θ′) (1).

[0117] Among them, y is the output target value, generally the set reference, that is, the d-axis current reference amount I dref and the q-axis current reference amount I qref ; θ represents the Q-value parameter of the DQN network. This parameter θ is the parameter of the corresponding function and is the parameter of Q′ in this expression. Since there are many parameters in the Q-value function during the training process of the reinforcement learning algorithm and this parameter is changing, the symbol θ′ is used to represent it. a is the expected output action amount. γ is a constant, set according to specific circumstances. It should be noted here that the quantity with the superscript ()′ represents the state quantity at the new moment. r represents the expected reward return, that is, the expected return value, which is a variable.

[0118] Step 126, update the network: Update the DQN network Q(s, a; θ) using the mean square error loss function:

[0119] L = E (s,a,r,s′) [(y - Q(s, a; θ)) 2 (2).

[0120] Among them, L is the abbreviation of LOSS, representing the loss function; E is the abbreviation of ERROR, representing the error between the actual value and the target value. Using formula (2), continuously calculate the difference between the Q value and the target value Q′ to update the DQN network, and at the same time update the experience tuple (s, a, r, s′). The purpose of updating the DQN network is to make the control action of the DQN network continuously adapt to the model. Because reinforcement learning requires a training process, through continuous training, the degree of model adaptation will be higher and the output control amount will be more accurate.

[0121] Step 127, update the target network: Periodically copy the parameters of the DQN network to the target network.

[0122] According to the initialization of the target network Q′(s, a; θ′) and the expression of the experience tuple (s t , a t , r t , s t+1 ), it can be seen that the parameters of the DQN network are r, Q′(s i+1, a′; θ′). It can be updated directly by copying because the target network is updated less frequently than the DQN network. Generally, the target network is updated only once after the DQN network has been updated multiple times.

[0123] Step 13. Termination condition: When certain conditions are met (such as the number of training times reaching the threshold or the performance index reaching the target), stop the training.

[0124] The control quantity output by the DQN network is a, and a includes the d-axis voltage reference V dref and the q-axis voltage reference V qref , which is the control signal. After training in Step 11, Step 13, and Step 13, the d-axis voltage reference V dref and the q-axis voltage reference V qref output by the network after training are more accurate.

[0125] According to Figure 10 and Figure 11 , it can be known that the agent of DQN reinforcement learning outputs the d-axis reference voltage V dref , the q-axis reference voltage V qref , Figure 10 In derror , the PI controller using the DQN reinforcement learning algorithm replaces the d-axis current error i qerror and the q-axis current error i dref , and outputs the d-axis reference voltage V qref , the q-axis reference voltage V ref . The agent of this DQN has six observed quantities, which are the speed reference ω fb , the speed feedback ω d , the direct-axis current i dref , the direct-axis current reference i q , the quadrature-axis current i qref . The expression of the expected return reward is set as shown in formula (3). The DQN network makes actions by maximizing the cumulative expected return in each episode.

[0126] The expression of the current reward r t is as follows:

[0127]

[0128] Formula (3) is the reward function of the DQN algorithm without the fuzzy logic controller. In formula (3), Q1 and Q2 are the Q-value functions of i derror and i qerror , that is, the action value function; i derror is the direct-axis current error; i qerror is the quadrature-axis current error; μ jis the action function parameter for the previous time step, j represents the number of actions, μ is also the action function of the DQN network; τ is a constant, usually taken from 0.05 to 0.2.

[0129] Based on Figure 12 、 Figure 13 、 Figure 14 and Table 1, substituting the obtained weight coefficient λ into formula (3), we can get:

[0130]

[0131] After obtaining the new reward function, the new target value can be calculated through the following formula:

[0132] y t = r t + γmax a′ Q′(s t+1 , a′; θ′) (6).

[0133] In formula (6), γ is a constant, Q′ represents the action value function of the target network; similarly, s′, a′, and θ′ represent the state quantity, action quantity, and Q-value function parameters of the target network respectively. Then calculate the difference between the DQN network target value and the target network evaluation value:

[0134] L c = E (s,a,r,s′) [(y t - Q(s, a; θ)) 2 (7).

[0135] Among them, L c Similarly, it also means the loss function. c is the abbreviation of critic, representing the error between the evaluation value. Update the parameters in the target network according to formula (7). When certain conditions are met (such as the number of training times reaches the threshold or the performance index reaches the target), stop training.

[0136] Figure 9 is Figure 10 the schematic diagram of the program loop training process of the DQN reinforcement learning algorithm in the reinforcement learning agent in dref and V qref are output as shown in the figure, saving the problem of separately adding the PI controller and tuning the PI controller parameters for each of the above two quantities. Figure 11 is Figure 10 the schematic diagram of the two-level three-phase inverter in Figure 10 to facilitate better understanding of a 、S b and S c output by the pulse width generator in Figure 12, Figure 13 , Figure 14 Overall, it constitutes Figure 10 the fuzzy logic controller in Figure 12 and Figure 13 are the input quantities of the fuzzy logic controller, Figure 14 is the output quantity of the fuzzy logic controller.

[0137] Figure 15 is the flow schematic diagram of the motor drive control method for an air-conditioning compressor with fuzzy logic optimized DQN reinforcement learning field-oriented control. As Figure 15 shown, the motor drive control method for an air-conditioning compressor with fuzzy logic optimized DQN reinforcement learning field-oriented control includes:

[0138] Step S1, obtain the rotor angular velocity error ω r_error of the compressor and the rotor flux linkage error ψ r_error , and input the rotor angular velocity error ω r_error and the rotor flux linkage error ψ r_error into the fuzzy logic controller to obtain the weight coefficient λ. Among them, the fuzzy logic controller is the overall composed of Figure 12 , Figure 13 and Figure 14 and Table 1, in order to output the reward function weight coefficient λ.

[0139] Step S2, collect the direct-axis current sampling quantity (i.e., the d-axis current sampling quantity I d ) and the quadrature-axis current sampling quantity (i.e., the q-axis current sampling quantity I q ) of the compressor, and obtain the direct-axis current reference quantity (i.e., the d-axis current reference quantity I dref ) and the quadrature-axis current reference quantity (i.e., the q-axis current reference quantity I qref ) through the PI controller.

[0140] Step S3, input the weight coefficient λ, the direct-axis current sampling quantity (i.e., the d-axis current sampling quantity I d ), the quadrature-axis current sampling quantity (i.e., the q-axis current sampling quantity I q ), the direct-axis current reference quantity (i.e., the d-axis current reference quantity I dref ) and the quadrature-axis current reference quantity (i.e., the q-axis current reference quantity I qref ) into the agent module of the DQN algorithm, and output the direct-axis voltage reference quantity (i.e., the d-axis voltage reference quantity V dref ) and the quadrature-axis voltage reference quantity (i.e., the q-axis voltage reference quantity V qref ) by the agent module.

[0141] Step S4, for the direct-axis voltage reference quantity (i.e., the d-axis voltage reference quantity V dref) and quadrature-axis voltage reference (i.e., q-axis voltage reference V qref ) are sequentially subjected to transformation processing and pulse-width modulation processing to obtain three-phase control signals (such as the control signals S a 、S b and S c ) of a two-level three-phase inverter.

[0142] Step S5: Control the two-level three-phase inverter by using the three-phase control signals.

[0143] The solution of the present invention relates to the field of motor drive control in an air-conditioning system, and particularly to a motor drive control method for an air-conditioning compressor based on field-oriented control optimized by fuzzy logic and DQN reinforcement learning. In the field-oriented control in the related solution, the direct-axis current I d (i.e., d-axis current) and the quadrature-axis current I q (i.e., q-axis current) are adjusted based on a PI controller according to errors. Since the parameter tuning of the PI controller is mostly based on rectification experience, there are problems such as difficult tuning and long tuning time. The reinforcement learning control method proposed in the solution of the present invention is a control method based on a neural network, and has the ability to automatically tune the corresponding weight coefficients and bias amounts by using the backpropagation algorithm according to parameters; moreover, the control method optimized based on a fuzzy logic controller proposed in the solution of the present invention can further decouple the motor parameters to achieve the purpose of better controlling the motor.

[0144] Adopting the technical solution of this embodiment, for the motor drive control method (such as the motor drive control method in an air-conditioning compressor) and the inverter used thereby (such as a two-level three-phase inverter), sample the three-phase current of the motor (such as the three-phase current I a 、I b and I c ) output by the two-level three-phase inverter, sample the actual speed of the motor (such as the feedback sampling amount ω fb ) of the rotor angular velocity of the motor, and obtain the target speed of the motor (such as the reference amount ω ref ) of the given rotor angular velocity of the motor, and the target amount of the rotor magnetic flux of the motor (such as the reference amount |ψ r_ref |) of the rotor magnetic flux; according to the three-phase current of the motor (such as the three-phase current I a 、I b and I c ) output by the two-level three-phase inverter, obtain the d-axis current sampling amount I d and the q-axis current sampling amount I q of the motor, and the sampling amount of the rotor magnetic flux of the motor (sampling amount |ψ r |) through the processing of an estimator and a converter; the actual speed of the motor (such as the feedback sampling amount ωf b) The absolute value of the rotor speed difference from the target speed of the motor (such as the given rotor angular velocity reference ω of the motor) ref ) (i.e., the rotor angular velocity error |ω r_error |) passes through the speed loop PI controller to obtain the q-axis current reference I qref . The sampled rotor flux of the motor (the sampled rotor flux |ψ r |), the absolute value of the difference between the target rotor flux of the motor (such as the rotor flux reference |ψ r_ref |) (i.e., the rotor flux error |ψ r_error |) passes through the flux loop PI controller to obtain the d-axis current reference I dref . The absolute value of the rotor speed difference (i.e., the rotor angular velocity error |ω r_error |), the absolute value of the rotor flux difference (i.e., the rotor flux error |ψ r_error |) pass through the fuzzy logic controller to obtain the weight coefficient λ; the weight coefficient λ, the q-axis current reference I qref , the sampled q-axis current I q , the d-axis current reference I dref , the sampled d-axis current I d , after being processed by DQN reinforcement learning, obtain the d-axis voltage reference V dref and the q-axis voltage reference V qref ; the d-axis voltage reference V dref and the q-axis voltage reference V qref , through the converter and pulse width modulation processing, obtain the control signals of each switching tube in the inverter to control the inverter to achieve the drive control of the motor; thus, in the motor drive control, based on the three-phase current and speed of the motor, by adopting the FOC control combining fuzzy logic processing and DQN reinforcement learning processing, control the inverter to achieve the drive control of the motor and improve the control effect of the motor.

[0145] According to an embodiment of the present invention, there is also provided a control device for a motor corresponding to the control method of the motor. Refer to Figure 7 the structural schematic diagram of an embodiment of the device of the present invention shown. The drive control end of the motor has an inversion module (such as a two-level three-phase inverter), and the inversion module has switching tubes; the inversion module is used to perform inversion processing on the input DC voltage source under the control of the control signals of the switching tubes to drive and control the motor. As Figure 11 shown, the main circuit of the two-level three-phase inverter circuit includes: the DC bus voltage V dc , the DC bus capacitor composed of capacitor C1 and capacitor C2 connected in series, the switching tube S a , the switching tube S b , the switching tube S c , the switching tube switching tube Switching tube A three-phase inverter composed of, and a permanent magnet synchronous motor (PMSM).

[0146] Among them, the DC bus voltage V dc The positive pole is connected to the positive pole of capacitor C1, and the negative pole of capacitor C1 is connected to the positive pole of capacitor C2; the DC bus voltage V dc The negative pole is connected to the negative pole of capacitor C2, and the midpoint of capacitor C1 and capacitor C2 is point o. The positive pole of the DC bus voltage V dc is also respectively connected to the first connection ends of each switching tube in switching tube S a , switching tube S b , switching tube S c ; the negative pole of the DC bus voltage V dc is also respectively connected to the second connection ends of each switching tube in switching tube Switching tube Switching tube ; the second connection ends of each switching tube in switching tube S a , switching tube S b , switching tube S c are correspondingly connected to the first connection ends of each switching tube in switching tube Switching tube Switching tube . The common end of the second connection end of switching tube S a and the first connection end of switching tube is point a, and point a is connected to the first connection end of the first-phase winding in the three-phase windings of the permanent magnet synchronous motor (PMSM); the common end of the second connection end of switching tube S b and the first connection end of switching tube is point b, and point b is connected to the first connection end of the second-phase winding in the three-phase windings of the permanent magnet synchronous motor (PMSM); the common end of the second connection end of switching tube S c and the first connection end of switching tube is point c, and point c is connected to the first connection end of the third-phase winding in the three-phase windings of the permanent magnet synchronous motor (PMSM); the second connection end of the first-phase winding, the second connection end of the second-phase winding, and the third connection end of the third-phase winding are all connected to point n. In the solution of the present invention, as Figure 7 shown, the control device of the motor includes: an acquisition unit 102 and a control unit 104.

[0147] Among them, the acquisition unit 102 is configured to, when the motor is running, acquire the three-phase current of the motor, acquire the actual speed of the motor, acquire the target speed of the motor, and acquire the target value of the rotor magnetic flux of the motor; among them, the three-phase current of the motor is like the three-phase current I a , I band I c wherein the actual rotational speed of the motor, such as the feedback sampling amount ωf of the angular velocity of the motor rotor b wherein the target rotational speed of the motor, such as the given reference amount ω of the angular velocity of the motor rotor ref wherein the target amount of the rotor magnetic flux of the motor, such as the reference amount of the rotor magnetic flux |ψ r_ref |. For the specific functions and processing of the acquisition unit 102, refer to step S110.

[0148] The control unit 104 is configured to obtain the control signals of the switching tubes in the inverter module by using FOC control combining fuzzy logic processing and DQN reinforcement learning processing according to the three-phase current of the motor, the actual rotational speed of the motor, the target rotational speed of the motor, and the target amount of the rotor magnetic flux of the motor; specifically, according to the three-phase current of the motor, the actual rotational speed of the motor, the target rotational speed of the motor, and the target amount of the rotor magnetic flux of the motor, after being estimated by an estimator, transformed by a converter, PI controlled, fuzzy logic controlled, DQN reinforcement learning processed, and pulse width modulated, the control signals of the switching tubes in the inverter module are obtained. Among them, the converter is used to perform Park transformation and inverse Park transformation. PI control is the control using a PI controller, fuzzy logic control is the control using a fuzzy logic controller, DQN reinforcement learning processing is the processing using a DQN reinforcement learning algorithm, and pulse width modulation processing is the processing using a pulse width generator. For the specific functions and processing of the control unit 104, refer to step S120.

[0149] The control unit 104 is further configured to control the operation of the inverter module according to the control signals of the switching tubes in the inverter module to achieve the control of the motor, specifically to achieve the drive control of the motor. For the specific functions and processing of the control unit 104, refer to step S130.

[0150] For the DQN reinforcement learning algorithm proposed by the solution of the present invention, the reinforcement learning algorithm has the ability to handle complex operation and complex control problems, and can significantly improve the energy efficiency, stability and performance of the system; the DQN reinforcement learning algorithm interacts with the environment and adjusts the decisive parameters based on the expected return, and finally achieves the role of target control by outputting the most suitable action parameters. DQN can dynamically adjust the working mode of the three-phase inverter according to the real-time load conditions and power supply characteristics to achieve a higher energy efficiency ratio.

[0151] In some embodiments, the control unit 104 obtains the control signals of the switching tubes in the inverter module by using FOC control combining fuzzy logic processing and DQN reinforcement learning processing according to the three-phase current of the motor, the actual rotational speed of the motor, the target rotational speed of the motor, and the target amount of the rotor magnetic flux of the motor, including:

[0152] The control unit 104 is specifically further configured to, based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target rotor magnetic flux of the motor, after being estimated by an estimator, transformed by a converter, and subjected to PI control by a PI controller, obtain the weight coefficient of the reward function in the DQN reinforcement learning process, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor; wherein, the weight coefficient of the reward function is such as the weight coefficient λ, and the sampled d-axis current of the motor is such as the sampled d-axis current I d , and the sampled q-axis current of the motor is such as the sampled q-axis current I q , and the reference d-axis current of the motor is such as the reference d-axis current I dref , and the reference q-axis current of the motor is such as the reference q-axis current I qref . For the specific functions and processes of this control unit 104, refer to step S210

[0153] The control unit 104 is specifically further configured to, based on the weight coefficient of the reward function, the sampled d-axis current of the motor, the sampled q-axis current of the motor, the reference d-axis current of the motor, and the reference q-axis current of the motor, after being subjected to DQN reinforcement learning processing by the DQN reinforcement learning algorithm, obtain the reference d-axis voltage of the motor and the reference q-axis voltage of the motor, and the reference d-axis voltage of the motor is such as the reference d-axis voltage V dref , and the reference q-axis voltage of the motor is such as the reference q-axis voltage V qref . For the specific functions and processes of this control unit 104, refer to step S220

[0154] The control unit 104 is specifically further configured to, based on the reference d-axis voltage of the motor and the reference q-axis voltage of the motor, after performing inverse Park transformation by a converter and pulse width generation processing by a pulse width generator, obtain the control signals of the switching tubes in the inverter module. For the specific functions and processes of this control unit 104, refer to step S230

[0155] A motor drive control scheme for an air conditioner compressor based on DQN reinforcement learning magnetic field orientation control optimized by fuzzy logic proposed by the solution of the present invention uses the DQN reinforcement learning algorithm to replace the PI controllers on the d-axis and q-axis in the FOC algorithm, retains the PI controllers of the speed loop and the magnetic flux loop, and then, according to the rotor magnetic flux error amount |ψ r_error | and the rotor angular velocity error amount |ω r_error |, quickly locks the corresponding weight coefficient by using fuzzy logic judgment, and applies the weight coefficient λ to the previous reward function rt ( ) Obtain a new reward function to achieve decoupled control of the rotor magnetic flux ψr and the rotational speed (such as the rotor angular velocity ωf b ), making the calculation more convenient and the control more efficient.

[0156] In some embodiments, the control unit 104, based on the three-phase current of the motor, the actual rotational speed of the motor, the target rotational speed of the motor, and the target value of the rotor magnetic flux of the motor, after being estimated by an estimator, transformed by a converter, and subjected to PI control by a PI controller, obtains the weight coefficient of the reward function in the DQN reinforcement learning process, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor, including:

[0157] The control unit 104 is specifically further configured to, based on the three-phase current of the motor, be estimated by an estimator and subjected to Park transformation by a converter to obtain the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, and the sampled value of the rotor magnetic flux of the motor; wherein, the sampled value of the d-axis current of the motor is such as the sampled value of the d-axis current I d , the sampled value of the q-axis current of the motor is such as the sampled value of the q-axis current I q , and the sampled value of the rotor magnetic flux of the motor is such as the sampled value of the rotor magnetic flux |ψ r |. For the specific functions and processes of this control unit 104, refer to step S310.

[0158] The control unit 104 is specifically further configured to determine the absolute value of the difference between the actual rotational speed of the motor and the target rotational speed of the motor, denoted as the rotational speed error of the motor, such as the angular velocity error |ω r_error |; and determine the absolute value of the difference between the sampled value of the rotor magnetic flux of the motor and the target value of the rotor magnetic flux of the motor, denoted as the rotor magnetic flux error of the motor, such as the rotor magnetic flux error |ψ r_error |. For the specific functions and processes of this control unit 104, refer to step S320.

[0159] The control unit 104 is specifically further configured to, based on the rotational speed error of the motor, after being subjected to PI control by a speed-loop PI controller, obtain the reference value of the q-axis current of the motor, such as the reference value of the q-axis current I qref ; based on the rotor magnetic flux error of the motor, after being subjected to PI control by a magnetic-flux-loop PI controller, obtain the reference value of the d-axis current of the motor, such as the reference value of the d-axis current I dref ; and based on the rotational speed error of the motor and the rotor magnetic flux error of the motor, after being subjected to fuzzy logic control by a fuzzy logic controller, obtain the weight coefficient of the reward function, specifically obtain the reward function r in the DQN reinforcement learning tThe weight coefficient, such as the weight coefficient λ. For the specific functions and processes of the control unit 104, refer to step S330.

[0160] In the solution of the present invention, according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target rotor flux of the motor, the PI controllers of the speed loop and the flux loop in the FOC control are retained. Furthermore, according to the rotor flux error amount |ψ r_error | and the rotor angular velocity error amount |ω r_error |, the fuzzy logic judgment is used to quickly lock the corresponding weight coefficient, and the weight coefficient λ is applied to the previous reward function r t () to obtain a new reward function, realizing the decoupling control of the rotor flux ψ r and the speed (such as the rotor angular velocity ω fb ), making the calculation more convenient and the control more efficient.

[0161] In some embodiments, the control unit 104 is further configured to pre-set fuzzy control logic in the fuzzy logic controller; the fuzzy control logic includes: input logic and output logic; wherein, the input logic includes: setting the correspondence between the speed error, the rotor flux error, and the set input parameter error range, such as Figure 12 、 Figure 13 and the relationship reflected in Table 1; the output logic includes: setting the correspondence between the weight coefficient and the set output parameter range, such as Figure 14 reflected relationship.

[0162] The control unit 104 obtains the weight coefficient of the reward function after performing fuzzy logic control through the fuzzy logic controller according to the speed error of the motor and the rotor flux error of the motor, including:

[0163] The control unit 104 is specifically further configured to determine the input parameter error range corresponding to the speed error of the motor and the rotor flux error of the motor by using the correspondence between the set speed error, the set rotor flux error, and the set input parameter error range according to the speed error of the motor and the rotor flux error of the motor. For the specific functions and processes of the control unit 104, refer to step S410.

[0164] The control unit 104 is specifically further configured to use a fuzzy interference system and the correspondence between a set weight coefficient and a set output parameter range to determine a weight coefficient corresponding to the rotational speed error of the motor and the rotor flux linkage error of the motor from the input parameter error range corresponding to the rotational speed error of the motor and the rotor flux linkage error of the motor, and use this as the weight coefficient of the reward function. For the specific functions and processing of this control unit 104, refer to step S420 for details.

[0165] Figure 12 is a schematic curve diagram of the input function of the fuzzy logic controller (such as the relationship between the rotor flux linkage error |ψ r_error | and VS, S, M, L, VL). Figure 13 is a schematic curve diagram of the input function of the fuzzy logic controller (such as the relationship between the rotor angular velocity error |ω r_error | and VS, S, M, L, VL). In addition, the input of the fuzzy logic needs to select the most suitable input parameter error range according to the rotor flux linkage error quantity such as the rotor angular velocity error |ω r_error | and the rotor angular velocity error quantity such as the rotor flux linkage error |ψ r_error |. Here, the rotor flux linkage error range is set to [0, 0.001], with the unit of wb (Weber), as Figure 12 shown; the rotor angular velocity error range is [0, 3.5], with the unit of rad / s (radian per second), as Figure 13 shown. That is, when each error falls within the above range, the corresponding weight coefficient λ can be quickly determined through the fuzzy interference system (Fuzzy Inference System, FIS). It should be noted here that when the rotor flux linkage error is less than 0 or the current error is greater than 0.001, the input amplitude needs to be limited within the range of [0, 0.001]. The same applies to the rotor angular velocity error range. That is:

[0166]

[0167] Figure 14 is a schematic curve diagram of the output function of the fuzzy logic controller (such as the relationship between λ and NL, NM, NS, O, PS, PM, PL). In the control method proposed by the solution of the present invention, a fuzzy logic controller is used to achieve the decoupling of the rotor flux linkage ψ r and the rotor angular velocity ω fb . The above two parameters, that is, the error quantities of the rotor angular velocity error |ω r_error | and the rotor flux linkage error |ψ r_error |, are used as the input of the fuzzy logic controller. After passing through the FIS, a weight coefficient λ is output. The range of this coefficient is 0 < λ ≤ 1, as Figure 14 shown. The principle for selecting the weight coefficient λ is shown in Table 1.

[0168] Figure 14 Each parameter in represents the value range of the weight coefficient λ, which is obtained by fuzzy judgment based on the flux linkage error and angular velocity error collected in Table 1. The parameter that the fuzzy logic controller needs to obtain is Figure 14 the weight coefficient λ in.

[0169] Table 1 and Figure 12 、 Figure 13 The horizontal axis in is the absolute value of the error amount (hereinafter simply referred to as the error), and VS, S, M, L, VL represent the fuzzy judgment values when the error is located in different ranges. VS = Very Small (very small), S = Small (small), M = medium (medium), L = Large (large), VL = Very Large (very large). The character meanings in the table are: L is the abbreviation of Large (large), M is the abbreviation of Medium (medium), S is the abbreviation of Small (small), P is the abbreviation of Positive (positive), and N is the abbreviation of Negative (negative).

[0170] In related solutions, when using FOC control, the quantities output by the motor are collected, such as the rotor flux linkage and rotor speed. Due to a certain degree of coupling between the above parameters, there is an interference problem during control. Therefore, the solution of the present invention proposes a fuzzy logic controller, which can complete the decoupling control of the motor parameters by fuzzyifying the rotor flux linkage and speed. Aiming at the problem of strong coupling of motor parameters, the solution of the present invention uses a fuzzy logic controller to achieve a certain degree of decoupling of the motor speed (such as the rotor angular velocity ω fb ) and the rotor flux linkage ψ r , realizes the decoupling of the rotor flux linkage error and the rotor angular velocity error, and achieves the purpose of better controlling the motor.

[0171] In some embodiments, the control unit 104, according to the weight coefficient of the reward function, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor, after performing DQN reinforcement learning processing through the DQN reinforcement learning algorithm, obtains the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor, including:

[0172] The control unit 104 is specifically further configured to use the DQN reinforcement learning algorithm to pre-train a DQN reinforcement learning model, such as the agent module of the DQN algorithm. For the specific functions and processes of this control unit 104, refer to step S510.

[0173] The control unit 104 is further specifically configured to use the DQN reinforcement learning model to obtain a d-axis voltage reference quantity and a q-axis voltage reference quantity of the motor according to the weight coefficient of the reward function, the sampled d-axis current quantity of the motor, the sampled q-axis current quantity of the motor, the d-axis current reference quantity of the motor, and the q-axis current reference quantity of the motor. For the specific functions and processes of this control unit 104, refer to step S520.

[0174] Figure 10 The overall system schematic diagram of the FOC control PMSM based on DQN reinforcement learning optimized by fuzzy logic. In the related solutions, for the FOC algorithm, two controllers (i.e., two PI controllers) are separately used for the d-axis control loop and the q-axis control loop; Figure 10 In the example shown, the DQN reinforcement learning algorithm is used to replace the PI controllers of the d-axis and q-axis in the FOC algorithm, and the PI controllers of the speed loop and the flux linkage loop are retained. Figure 10 In the example shown, it includes a speed loop PI controller and a flux linkage loop PI controller.

[0175] Figure 10 where, |ω r_error | represents the rotor angular velocity error; |ψ r_error | represents the rotor flux linkage error; ω ref represents the rotor angular velocity reference quantity; I qref represents the q-axis current reference quantity; I q represents the sampled q-axis current quantity; |ψ r | represents the sampled rotor flux linkage quantity; |ψ r_ref | represents the rotor flux linkage reference quantity; I dref represents the d-axis current reference quantity; I d represents the sampled d-axis current quantity; V qref represents the q-axis voltage reference quantity; V dref represents the d-axis voltage reference quantity; V αref represents the α-axis voltage reference quantity; V βref represents the β-axis voltage reference quantity; S a , S b , S c represents the a, b, c phase switch control quantity; V dc represents the DC voltage source voltage; I a represents the sampled a-phase current quantity; I b represents the sampled b-phase current quantity; I α represents the sampled α-axis current quantity; I β represents the sampled β-axis current quantity; θ e represents the converter a-phase angle quantity; ω fb represents the sampled rotor angular velocity feedback quantity.

[0176] The overall block diagram of the FOC control PMSM with fuzzy logic DQN is as follows Figure 10 As shown, according to Figure 10 it can be seen that among the three-phase currents I a , I b and I c output by the two-level three-phase inverter, the sampled value of the a-phase current I a and the sampled value of the b-phase current I b , through the estimator, the sampled value of the α-axis current I α , the sampled value of the β-axis current I β , the converter a-phase angle measurement θ e , and the rotor flux sampled value |ψ r | are obtained. Based on the converter a-phase angle measurement θ e , after the sampled value of the α-axis current I α and the sampled value of the β-axis current I β are subjected to Park transformation, the sampled value of the d-axis current I d and the sampled value of the q-axis current I q are output. An estimator is a core tool in statistics and machine learning for parameter estimation or model construction, which infers the overall characteristics or model parameters through sample data. The d-axis current reference I dref is obtained by passing the rotor flux error |ψ r | obtained by comparing the rotor flux sampled value |ψ r_ref | with the rotor flux reference |ψ r_error | through a PI controller. The q-axis current reference I qref is obtained by passing the rotor angular velocity error |ω fb | obtained by comparing the rotor angular velocity feedback sampled value ω ref with the given rotor angular velocity reference ω r_error | through a PI controller. Park transformation (i.e., Park transformation) is to project physical quantities such as currents and voltages in the three-phase stationary coordinate system (abc) onto the rotating coordinate system (dq0) composed of the direct axis (d-axis), quadrature axis (q-axis), and zero axis (0-axis) that rotates with the rotor, thereby simplifying the dynamic model of the motor. The inverse Park transformation is a mathematical method for converting variables in the two-phase rotating coordinate system (dq coordinate system) back to the two-phase stationary coordinate system (αβ coordinate system), mainly used in motor control and power system analysis.

[0177] Among them, the rotor angular velocity error |ω r_error | and the rotor flux error |ψ r_error | are respectively used as the inputs of the fuzzy logic controller. The fuzzy logic controller outputs a weight coefficient λ to the DQN agent (i.e., the DQN reinforcement learning agent), and then through the DQN agent, the d-axis voltage reference V dref and the q-axis voltage reference V qref, the reference quantity (i.e., the d-axis voltage reference quantity V dref and the q-axis voltage reference quantity V qref ) After the inverse Park transformation, and then through the pulse width generation module, the control signals S a , S b and S c for controlling the switching tubes in the two-level three-phase inverter can be generated. The two-level three-phase inverter outputs three-phase voltages V a , V b and V c to the motor PMSM.

[0178] In related solutions, when using FOC to control a three-phase inverter, the d-axis control loop and the q-axis control loop require two PI controllers; there are problems with difficult parameter tuning for the two PI controllers under various different power supplies and loads. However, in the solution of the present invention, after applying the reinforcement learning technology, it is not necessary to complete the parameter tuning of the above-mentioned control loops, and the self-learning technology can be used to complete the adaptation and reference parameter output under specific circumstances. Aiming at the problem of difficult parameter tuning for the PI of the d-axis and q-axis controllers, the solution of the present invention uses the DQN reinforcement learning method to automatically tune the corresponding weight coefficients and bias amounts according to the parameters using the backpropagation algorithm, which is dynamically better and has the ability of self-learning and self-correction.

[0179] In some embodiments, the control unit 104 uses the DQN reinforcement learning algorithm to pre-train a DQN reinforcement learning model, including:

[0180] The control unit 104 is specifically further configured to collect the historical data of the motor as sample data; the historical data of the motor includes: the weight coefficients of the reward function, and the historical values of the d-axis current sampling amount, q-axis current sampling amount, d-axis current reference amount, q-axis current reference amount, d-axis voltage reference amount and q-axis voltage reference amount of the motor. For the specific functions and processing of this control unit 104, please also refer to step S610.

[0181] The control unit 104 is specifically further configured to, based on the DQN network, use the weight coefficients of the reward function, and the historical values of the d-axis current sampling amount, q-axis current sampling amount, d-axis current reference amount and q-axis current reference amount of the motor as input quantities, and use the historical values of the d-axis voltage reference amount and q-axis voltage reference amount of the motor as output quantities, and perform reinforcement learning training using the sample data to obtain the DQN reinforcement learning model. For the specific functions and processing of this control unit 104, please also refer to step S620.

[0182] Figure 8 is a schematic block diagram of the principle of the DQN algorithm. In Figure 8In the experience replay pool shown, it contains the content of two steps, one is storage and the other is replay. Storage means storing data such as experience tuples (s t , a t , r t+1 , s t+1 ) into the experience replay pool as experience; replay means collecting experience data from the experience replay pool through certain specific rules during training.

[0183] Figure 9 is a schematic flow diagram of the loop training process of the DQN algorithm. Figure 8 and Figure 9 in the environment represent the scenarios to which the DQN algorithm is applied. For example, the application scenario in the solution of the present invention is the motor drive control scenario. See Figure 8 , Figure 9 and Figure 10 for the examples shown. In the DQN algorithm, the state s contains five quantities input to the DQN reinforcement learning agent, namely (λ, I qref , I q , I dref , I d ). s0 represents the state at the initial moment, and generally the above values are set to 0. a is the output action. According to Figure 10 , it can be known that a = (V qref , V dref ). Iqref represents the q-axis current reference quantity, I q represents the q-axis current sampling quantity, I dref represents the d-axis current reference quantity, I d represents the d-axis current sampling quantity, and λ represents a constant weight coefficient.

[0184] In the solution of the present invention, the DQN reinforcement learning algorithm is used to achieve the role of target control by outputting the most appropriate action parameters, so as to dynamically adjust the working mode of the three-phase inverter according to the real-time load conditions and power supply characteristics to achieve a higher energy efficiency ratio.

[0185] In some embodiments, the control unit 104, based on the DQN network, uses the weight coefficient of the reward function, and the historical values of the sampled d-axis current, q-axis current, d-axis current reference, and q-axis current reference of the motor as input quantities, and the historical values of the d-axis voltage reference and q-axis voltage reference of the motor as output quantities, and performs reinforcement learning training using the sample data to obtain the DQN reinforcement learning model, including: the control unit 104 is specifically further configured to, based on the DQN network, use the weight coefficient of the reward function, and the historical values of the sampled d-axis current, q-axis current, d-axis current reference, and q-axis current reference of the motor as input quantities, and the historical values of the d-axis voltage reference and q-axis voltage reference of the motor as output quantities, and perform reinforcement learning training using the sample data according to the following formula to obtain the DQN reinforcement learning model:

[0186] y i = r i + γ max a′ Q′(s i+1 , a′; θ′);

[0187]

[0188] where i represents the i-th group of sampled quantities in the sample data; y i represents the output target value, representing the historical values of the d-axis current reference and q-axis current reference; r i represents the reward return value expected for reinforcement learning training using the sample data, which is a variable, and uses the reward function of the DQN algorithm. The reward function of the DQN algorithm is like the function r t (), t represents the time step; γ is a constant; a represents the expected output action quantity, representing the historical values of the d-axis voltage reference and q-axis voltage reference of the motor, a′ represents the expected output action quantity at the new moment, representing the historical values of the d-axis voltage reference and q-axis voltage reference of the motor at the new moment; Q represents the initialized network of the DQN network, Q′ represents the initialized target network of the DQN network; θ′ represents the parameters of Q′; Q1 represents the Q-value function of the direct-axis current error i derror , Q2 represents the Q-value function of the quadrature-axis current error i qerror ; μ j is the action function parameter of the previous time step, j represents the number of actions, and τ is a constant.

[0189] As Figure 8 and Figure 9 shown, the process flow of the loop training process of the DQN algorithm includes:

[0190] Step 11. Initialization: Initialize the experience replay buffer D; Initialize the DQN network Q(s, a; θ); Initialize the target network Q′(s, a; θ′); Initialize the state s0.

[0191] Among them, for initialization, it only needs to assign an initial value to each variable before the start of the program. Initialization is a crucial programming step, and its purpose is to ensure that variables, objects, and systems are in the expected initial state before the program runs. The experience replay buffer D is a buffer in the experience replay pool.

[0192] Among them, t represents the time step. s represents the state, s t represents the state at the current moment t, s t+1 represents the state at the next moment t + 1; a represents the action, a t represents the action at the current moment t, a t+1 represents the action at the next moment t + 1; θ represents the parameters involved in the update and calculation process, such as the parameters of the initialized DQN network Q; θ′ represents the parameters of the initialized target network Q′ (i.e., Target Q). r represents the reward, r t represents the reward at the current t, r t+1 represents the reward at the next moment t + 1.

[0193] Step 12. Loop training:

[0194] For each time step t: The time step t is each calculation cycle of the DQN network. Generally speaking, the smaller t is set, the more accurate the output control amount is, and the corresponding calculation complexity is greater. Determine the time step t to complete the loop update of the DQN network.

[0195] Step 121. Select an action: Select the current action a t according to the current state s t and the current policy (such as the ε-greedy policy);

[0196] Step 122. Execute the action: Execute the current action a t , observe the next state s t+1 and the reward r t .

[0197] Step 123. Store the experience: Store the experience tuple (s t , a t , r t , s t+1 ) into the experience replay buffer D.

[0198] Step 124. Sample the experience: Randomly sample a batch of experience tuples (s i , a i , ri , s i+1 ). i represents the i-th group of sampling amounts.

[0199] Step 125, calculate the target value:

[0200] y i = r i + γmax a′ Q′(s i+1 , a′; θ′) (1).

[0201] Where y is the output target value, generally the set reference, that is, the d-axis current reference quantity I dref and the q-axis current reference quantity I qref ; θ represents the Q-value parameter of the DQN network. This parameter θ is the parameter of the corresponding function and is the parameter of Q′ in this expression. Since there are many parameters in the Q-value function during the training process of the reinforcement learning algorithm and this parameter is changing, the symbol θ′ is used to represent it. a is the expected output action amount. γ is a constant, set according to specific circumstances. It should be noted here that the state quantity with the superscript ()′ represents the state quantity at the new moment. r represents the expected reward return, that is, the expected return value, which is a variable.

[0202] Step 126, update the network: Use the mean squared error loss function to update the DQN network Q(s, a; θ):

[0203] L = E (s,a,r,s′) [(y - Q(s, a; θ)) 2 (2).

[0204] Where L is the abbreviation of LOSS, representing the loss function; E is the abbreviation of ERROR, representing the error between the actual value and the target value. Using formula (2), continuously calculate the difference between the Q value and the target value Q′ to update the DQN network, and at the same time update the experience tuple (s, a, r, s′). The purpose of updating the DQN network is to make the control action of the DQN network continuously adapt to the model. Because reinforcement learning requires a training process, through continuous training, the degree of model adaptation will be higher and the output control amount will be more accurate.

[0205] Step 127, update the target network: Periodically copy the parameters of the DQN network to the target network.

[0206] According to the initialization of the target network Q′(s, a; θ′) and the expression of the experience tuple (s t , a t , r t , s t+1 ), it can be seen that the parameters of the DQN network are r, Q′(s i+1, a′; θ′). It can be updated directly by copying because the target network is updated less frequently than the DQN network. Generally, the target network is updated only once after the DQN network has been updated multiple times.

[0207] Step 13. Termination condition: When certain conditions are met (such as the number of training times reaching the threshold or the performance index reaching the target), stop the training.

[0208] The control quantity output by the DQN network is a, and a includes the d-axis voltage reference V dref and the q-axis voltage reference V qref , which is the control signal. After training in Step 11, Step 13, and Step 13, the d-axis voltage reference V dref and the q-axis voltage reference V qref output by the network after training are more accurate.

[0209] According to Figure 10 and Figure 11 , it can be known that for the agent of DQN reinforcement learning, the output is the d-axis reference voltage V dref , the q-axis reference voltage V qref , Figure 10 In derror , the PI controller using the DQN reinforcement learning algorithm replaces the d-axis current error i qerror and the q-axis current error i dref , and outputs the d-axis reference voltage V qref , the q-axis reference voltage V ref . The six observables set for this DQN agent are the rotational speed reference ω fb , the rotational speed feedback ω d , the direct-axis current i dref , the direct-axis current reference i q , the quadrature-axis current i qref , and the quadrature-axis current reference i

[0210] The expression for the current reward r t is as follows:

[0211]

[0212] Equation (3) is the reward function of the DQN algorithm without a fuzzy logic controller. In Equation (3), Q1 and Q2 are the Q-value functions of i derror and i qerror , that is, the action value function; i derror is the direct-axis current error; i qerror is the quadrature-axis current error; μ jis the action function parameter of the previous time step, j represents the number of actions, and μ is also the action function of the DQN network; τ is a constant, usually taken from 0.05 to 0.2.

[0213] Based on Figure 12 、 Figure 13 、 Figure 14 and Table 1, substituting the obtained weight coefficient λ into formula (3), we can get:

[0214]

[0215] After obtaining the new reward function, the new target value can be calculated through the following formula:

[0216] y t = r t + γ max a′ Q′(s t+1 , a′; θ′) (6).

[0217] In formula (6), γ is a constant, and Q′ represents the action value function of the target network; similarly, s′, a′, and θ′ respectively represent the state quantity, action quantity, and Q-value function parameter of the target network. Furthermore, calculate the difference between the DQN network target value and the target network evaluation value:

[0218] L c = E (s,a,r,s′) [(y t - Q(s, a; θ)) 2 (7).

[0219] Update the parameters in the target network according to formula (7). When certain conditions are met (such as the number of training times reaching the threshold or the performance index reaching the target), stop training.

[0220] Figure 9 is Figure 10 the schematic diagram of the program loop training process of the DQN reinforcement learning algorithm in the reinforcement learning agent in dref and V qref , saving the problem of separately adding the PI controller and tuning the PI controller parameters for each of the above two quantities. Figure 11 is Figure 10 the schematic diagram of the two-level three-phase inverter in Figure 10 to facilitate a better understanding of a 、S b and S c in Figure 12 , Figure 13 , Figure 14 Generally speaking, it constitutes Figure 10The fuzzy logic controller in Figure 12 and Figure 13 are the input quantities of the fuzzy logic controller, Figure 14 is the output quantity of the fuzzy logic controller.

[0221] Figure 15 is a schematic flow diagram of the motor drive control method for an air-conditioning compressor with magnetic field orientation control based on DQN reinforcement learning optimized by fuzzy logic. As Figure 15 shown, the motor drive control method for an air-conditioning compressor with magnetic field orientation control based on DQN reinforcement learning optimized by fuzzy logic includes:

[0222] Step S1: Obtain the rotor angular velocity error |ω r_error | and the rotor magnetic flux error |ψ r_error | of the compressor, and input the rotor angular velocity error |ω r_error | and the rotor magnetic flux error |ψ r_error | into the fuzzy logic controller to obtain the weight coefficient λ.

[0223] Step S2: Collect the direct-axis current sampling quantity (i.e., the d-axis current sampling quantity I d ) and the quadrature-axis current sampling quantity (i.e., the q-axis current sampling quantity I q ) of the compressor, and obtain the direct-axis current reference quantity (i.e., the d-axis current reference quantity I dref ) and the quadrature-axis current reference quantity (i.e., the q-axis current reference quantity I qref ) through a PI controller.

[0224] Step S3: Input the weight coefficient λ, the direct-axis current sampling quantity (i.e., the d-axis current sampling quantity I d ), the quadrature-axis current sampling quantity (i.e., the q-axis current sampling quantity I q ), the direct-axis current reference quantity (i.e., the d-axis current reference quantity I dref ) and the quadrature-axis current reference quantity (i.e., the q-axis current reference quantity I qref ) into the agent module of the DQN algorithm, and output the direct-axis voltage reference quantity (i.e., the d-axis voltage reference quantity V dref ) and the quadrature-axis voltage reference quantity (i.e., the q-axis voltage reference quantity V qref ) by the agent module.

[0225] Step S4: Perform transformation processing and pulse width modulation processing on the direct-axis voltage reference quantity (i.e., the d-axis voltage reference quantity V dref ) and the quadrature-axis voltage reference quantity (i.e., the q-axis voltage reference quantity V qref ) in sequence to obtain three-phase control signals (such as the control signals S a , S b and S c ) of the switches in a two-level three-phase inverter.

[0226] Step S5: Use the three-phase control signal to control the two-level three-phase inverter.

[0227] The solution of the present invention relates to the field of motor drive control in an air-conditioning system, and particularly to a motor drive control method for an air-conditioning compressor based on field-oriented control optimized by fuzzy logic and DQN reinforcement learning. In the field-oriented control of the related solution, the direct-axis current I d (i.e., the d-axis current) and the quadrature-axis current I q (i.e., the q-axis current) are adjusted based on a PI controller for error. Since the parameter tuning of the PI controller is mostly based on rectification experience, there are problems such as difficult tuning and long tuning time. The reinforcement learning control method proposed in the solution of the present invention is a control method based on a neural network, which has the ability to automatically tune the corresponding weight coefficients and bias amounts using the backpropagation algorithm according to parameters; and the control method optimized based on a fuzzy logic controller proposed in the solution of the present invention can further decouple the motor parameters to achieve better motor control.

[0228] Since the processing and functions implemented by the device in this embodiment are basically corresponding to the embodiments, principles, and examples of the foregoing method, for the details not described in the description of this embodiment, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0229] According to an embodiment of the present invention, there is also provided a motor corresponding to the control device of the motor. The motor may include: the control device of the motor described above.

[0230] Since the processing and functions implemented by the motor in this embodiment are basically corresponding to the embodiments, principles, and examples of the foregoing device, for the details not described in the description of this embodiment, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0231] According to an embodiment of the present invention, there is also provided a compressor corresponding to the control device of the motor. The compressor may include: the control device of the motor described above, or the motor described above.

[0232] Since the processing and functions implemented by the compressor in this embodiment are basically corresponding to the embodiments, principles, and examples of the foregoing device, for the details not described in the description of this embodiment, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0233] According to an embodiment of the present invention, there is also provided an air conditioner corresponding to the control device of the motor. The air conditioner may include: the control device of the motor described above, or the motor described above, or the compressor described above.

[0234] Since the processing and functions implemented by the air conditioner in this embodiment are basically corresponding to the embodiments, principles, and examples of the foregoing device, for the details not described in the description of this embodiment, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0235] According to an embodiment of the present invention, there is also provided a computer program product corresponding to the control method of the motor, including a computer program, and when the computer program is executed by a processor, the steps of the control method of the motor described above are implemented.

[0236] Since the processing and functions implemented by the product in this embodiment are basically corresponding to the embodiments, principles, and examples of the foregoing method, for the details not described in the description of this embodiment, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0237] According to an embodiment of the present invention, there is also provided a storage medium corresponding to the control method of the motor, and the storage medium includes a stored program, wherein when the program runs, the device where the storage medium is located is controlled to execute the steps of the control method of the motor described above.

[0238] Since the processing and functions implemented by the storage medium in this embodiment are basically corresponding to the embodiments, principles, and examples of the foregoing method, for the details not described in the description of this embodiment, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0239] In summary, it is easy for those skilled in the art to understand that, on the premise of no conflict, the above-mentioned advantageous methods can be freely combined and superimposed.

[0240] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A control method for an electric motor, characterized in that, The drive control end of the motor has an inverter module, and the inverter module has switching tubes; the control method of the motor includes: Obtain the three-phase current of the motor, obtain the actual speed of the motor, obtain the target speed of the motor, and obtain the target value of the rotor magnetic flux of the motor; After performing FOC control combining fuzzy logic processing and DQN reinforcement learning processing based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, obtain the control signal of the switching tubes in the inverter module; Control the operation of the inverter module according to the control signal of the switching tubes in the inverter module to achieve the control of the motor.

2. The control method of the motor according to claim 1, characterized in that After performing FOC control combining fuzzy logic processing and DQN reinforcement learning processing based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, obtaining the control signal of the switching tubes in the inverter module includes: After estimating by an estimator, transforming by a converter, and performing PI control by a PI controller based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, obtain the weight coefficient of the reward function in the DQN reinforcement learning processing, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor; After performing DQN reinforcement learning processing by the DQN reinforcement learning algorithm based on the weight coefficient of the reward function, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor, obtain the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor; After performing inverse Park transformation by a converter and pulse width generation processing by a pulse width generator based on the reference value of the d-axis voltage of the motor and the reference value of the q-axis voltage of the motor, obtain the control signal of the switching tubes in the inverter module.

3. The control method of the motor according to claim 2, wherein After estimating by an estimator, transforming by a converter, and performing PI control by a PI controller based on the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, obtaining the weight coefficient of the reward function in the DQN reinforcement learning processing, the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, the reference value of the d-axis current of the motor, and the reference value of the q-axis current of the motor includes: After estimating by an estimator and performing Park transformation by a converter based on the three-phase current of the motor, obtain the sampled value of the d-axis current of the motor, the sampled value of the q-axis current of the motor, and the sampled value of the rotor magnetic flux of the motor; Determine the absolute value of the difference between the actual speed of the motor and the target speed of the motor, denoted as the speed error of the motor; and determine the absolute value of the difference between the sampled value of the rotor magnetic flux of the motor and the target value of the rotor magnetic flux of the motor, denoted as the rotor magnetic flux error of the motor. Based on the rotational speed error of the motor, after PI control by the speed loop PI controller, the q-axis current reference of the motor is obtained; based on the rotor flux linkage error of the motor, after PI control by the flux linkage loop PI controller, the d-axis current reference of the motor is obtained; and based on the rotational speed error of the motor and the rotor flux linkage error of the motor, after fuzzy logic control by the fuzzy logic controller, the weight coefficient of the reward function is obtained.

4. The control method of the motor according to claim 3, wherein, Fuzzy control logic is preset in the fuzzy logic controller; The fuzzy control logic includes: input logic and output logic; among them, the input logic includes: setting the correspondence relationship between the rotational speed error, the rotor flux linkage error, and the set input parameter error range; the output logic includes: setting the correspondence relationship between the weight coefficient and the set output parameter range; Obtaining the weight coefficient of the reward function after fuzzy logic control by the fuzzy logic controller according to the rotational speed error of the motor and the rotor flux linkage error of the motor includes: According to the rotational speed error of the motor and the rotor flux linkage error of the motor, using the correspondence relationship between the set rotational speed error, the set rotor flux linkage error, and the set input parameter error range, determine the input parameter error range corresponding to the rotational speed error of the motor and the rotor flux linkage error of the motor; Using the input parameter error range corresponding to the rotational speed error of the motor and the rotor flux linkage error of the motor, by using the fuzzy interference system and the correspondence relationship between the set weight coefficient and the set output parameter range, determine the weight coefficient corresponding to the rotational speed error of the motor and the rotor flux linkage error of the motor as the weight coefficient of the reward function.

5. The control method of the motor according to any one of claims 2 to 4, characterized in that, Obtaining the d-axis voltage reference of the motor and the q-axis voltage reference of the motor after DQN reinforcement learning processing by the DQN reinforcement learning algorithm according to the weight coefficient of the reward function, the d-axis current sampling value of the motor, the q-axis current sampling value of the motor, the d-axis current reference of the motor, and the q-axis current reference of the motor includes: Using the DQN reinforcement learning algorithm, pre-train to obtain a DQN reinforcement learning model; According to the weight coefficient of the reward function, the d-axis current sampling value of the motor, the q-axis current sampling value of the motor, the d-axis current reference of the motor, and the q-axis current reference of the motor, use the DQN reinforcement learning model to obtain the d-axis voltage reference of the motor and the q-axis voltage reference of the motor.

6. The control method of the motor according to claim 5, characterized in that, Using the DQN reinforcement learning algorithm to pre-train to obtain a DQN reinforcement learning model includes: Collect the historical data of the motor as sample data; the historical data of the motor includes: the weight coefficient of the reward function and the historical values of the d-axis current sampling value, q-axis current sampling value, d-axis current reference, q-axis current reference, d-axis voltage reference, and q-axis voltage reference of the motor; Based on the DQN network, using the weight coefficient of the reward function, as well as the historical values of the sampled d-axis current, sampled q-axis current, reference d-axis current, and reference q-axis current of the motor as input quantities, and using the historical values of the reference d-axis voltage and reference q-axis voltage of the motor as output quantities, perform reinforcement learning training using the sample data to obtain the DQN reinforcement learning model.

7. The control method of the motor according to claim 6, characterized in that, Based on the DQN network, using the weight coefficient of the reward function, as well as the historical values of the sampled d-axis current, sampled q-axis current, reference d-axis current, and reference q-axis current of the motor as input quantities, and using the historical values of the reference d-axis voltage and reference q-axis voltage of the motor as output quantities, perform reinforcement learning training using the sample data to obtain the DQN reinforcement learning model, including: Based on the DQN network, using the weight coefficient of the reward function, as well as the historical values of the sampled d-axis current, sampled q-axis current, reference d-axis current, and reference q-axis current of the motor as input quantities, and using the historical values of the reference d-axis voltage and reference q-axis voltage of the motor as output quantities, according to the following formula, perform reinforcement learning training using the sample data to obtain the DQN reinforcement learning model: y i = r i + γmax a′ Q′(s i+1 , a′; θ′); where i represents the i-th sampling amount in the sample data; y i represents the output target value, representing the historical values of the d-axis current reference quantity and the q-axis current reference quantity; r i represents the expected reward return value for reinforcement learning training using the sample data, which is a variable and uses the reward function of the DQN algorithm; γ is a constant; a′ represents the expected output action amount at the new moment, representing the historical values of the d-axis voltage reference quantity and the q-axis voltage reference quantity of the motor at the new moment; Q′ represents the initialized target network of the DQN network; θ′ represents the parameter of Q′.

8. A control device for an electric machine, characterized in that, The drive control end of the motor has an inverter module, and the inverter module has switching tubes; the control device of the motor includes: An acquisition unit configured to acquire the three-phase current of the motor, acquire the actual speed of the motor, acquire the target speed of the motor, and acquire the target value of the rotor magnetic flux of the motor. A control unit configured to, according to the three-phase current of the motor, the actual speed of the motor, the target speed of the motor, and the target value of the rotor magnetic flux of the motor, after performing FOC control combining fuzzy logic processing and DQN reinforcement learning processing, obtain the control signal of the switching tubes in the inverter module. The control unit is further configured to control the operation of the inverter module according to the control signal of the switching tubes in the inverter module to achieve the control of the motor.

9. A motor, characterized in that, Including: The control device of the motor as described in claim 8.

10. A compressor, characterized in that, Including: The control device of the motor as described in claim 8, or the motor as described in claim 9.

11. An air conditioner, characterized in that, Including: The control device of the motor as described in claim 8, or the motor as described in claim 9, or the compressor as described in claim 10.

12. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the control method of the motor described in any one of claims 1 to 7.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the control method of the motor described in any one of claims 1 to 7.