Joint motor control method and device, robot and storage medium
By updating network parameters by the reinforcement learning controller, generating interaxial and direct-axis control voltages, the problem of inaccurate motor control of robot joints is solved, and the complexity and accuracy of robot movements is improved.
Patent Information
- Application Number
- CN202410045406.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-11
AI Technical Summary
The robot's joint motor has poor accuracy during control, resulting in the inability to complete more complex actions.
The reinforcement learning controller is adopted to determine the intersection axis and the direct axis control current, combine the joint motor feedback current, and use the reinforcement learning model to update the network parameters, generate the intersection axis and the direct axis control voltage, and control the joint motor movement.
It improves the control accuracy of joint motors, improves the complexity and control accuracy of robot movements.
Smart Images

Figure CN120301262A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of robot technology, and in particular, to a method and device for controlling a joint motor, a robot, and a storage medium. Background Art
[0002] In recent years, the technology of robot technology has been continuously developing, becoming more and more intelligent and automated, and the richness, stability, and flexibility of movements have all been improved to varying degrees. A bionic robot has multiple joints, and a joint motor is provided at each joint. The joint motor can drive the parts on both sides of the joint to perform relative movement. It is precisely by the movement of these joint motors that the robot can complete various actions. However, in related technologies, the accuracy of controlling the joint motors of the robot is poor, resulting in the robot being unable to complete relatively complex actions. Summary of the Invention
[0003] To overcome the problems existing in related technologies, embodiments of the present disclosure provide a method and device for controlling a joint motor, a robot, and a storage medium to solve the defects in related technologies.
[0004] According to a first aspect of an embodiment of the present disclosure, a method for controlling a joint motor is provided. The method includes:
[0005] Determine a quadrature-axis control current according to a desired joint torque;
[0006] Input the quadrature-axis control current, the direct-axis control current, the quadrature-axis actual current, and the direct-axis actual current fed back by the joint motor into a reinforcement learning controller to obtain a quadrature-axis control voltage and a direct-axis control voltage output by the reinforcement learning controller;
[0007] Control the joint motor to move according to the quadrature-axis control voltage and the direct-axis control voltage;
[0008] Wherein, the reinforcement learning controller is configured to execute in each control frame:
[0009] Determine a reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the quadrature-axis actual current and the direct-axis actual current of the current control frame, and add an experience point of the previous control frame to the experience pool;
[0010] Update the network parameters of the reinforcement learning model according to a parameter set of at least one experience point in the experience pool, where the parameter set includes the state of the current control frame, the quadrature-axis control voltage, the direct-axis control voltage, the reward value, and the state of the next control frame, and the state includes the quadrature-axis actual current and the direct-axis actual current;
[0011] Generate a quadrature-axis control voltage and a direct-axis control voltage of the current control frame by using the reinforcement learning model.
[0012] In one embodiment of the present disclosure, determining the quadrature-axis control current according to the desired joint torque includes:
[0013] Determining the quadrature-axis control current according to the desired joint torque, the permanent magnet flux linkage amplitude of the joint motor, and the number of pole pairs.
[0014] In one embodiment of the present disclosure, the state further includes a quadrature-axis current error, a direct-axis current error, a quadrature-axis cumulative error, and a direct-axis cumulative error;
[0015] Wherein, the quadrature-axis current error includes the error between the quadrature-axis control current of the previous control frame and the quadrature-axis actual current of the current control frame;
[0016] The direct-axis current error includes the error between the direct-axis control current of the previous control frame and the direct-axis actual current of the current control frame;
[0017] The quadrature-axis cumulative error includes the sum of the quadrature-axis current errors of multiple control frames including the current control frame;
[0018] The direct-axis cumulative error includes the sum of the direct-axis current errors of multiple control frames including the current control frame.
[0019] In one embodiment of the present disclosure, when the reinforcement learning controller is used to determine the reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the quadrature-axis actual current and the direct-axis actual current of the current control frame, it is used for:
[0020] Determining the reward value according to the quadrature-axis current error and the direct-axis current error of the current control frame, and the quadrature-axis control voltage and the direct-axis control voltage of the previous frame.
[0021] In one embodiment of the present disclosure, the reinforcement learning controller is further used in each control frame:
[0022] Smoothing the network parameters of the reinforcement learning model according to the network parameters before and after the update of the reinforcement learning model.
[0023] In one embodiment of the present disclosure, when the reinforcement learning controller is used to generate the quadrature-axis control voltage and the direct-axis control voltage using the reinforcement learning model in each control frame, it is used for any one of the following:
[0024] Before the network parameters of the reinforcement learning model are updated, using the reinforcement learning model to generate the quadrature-axis control voltage and the direct-axis control voltage;
[0025] After the network parameters of the reinforcement learning model are updated, using the reinforcement learning model to generate the quadrature-axis control voltage and the direct-axis control voltage;
[0026] After the network parameters of the reinforcement learning model are smoothed, the reinforcement learning model is used to generate a quadrature-axis control voltage and a direct-axis control voltage.
[0027] In one embodiment of the present disclosure, the method further includes:
[0028] The desired joint state in the motion control instruction and the actual joint state feedback by the joint motor are input into an impedance controller to obtain the desired joint torque output by the impedance controller.
[0029] In one embodiment of the present disclosure, the desired joint state includes a desired joint position and a desired joint velocity; and / or,
[0030] The actual joint state includes an actual joint position and an actual joint velocity.
[0031] In one embodiment of the present disclosure, the method further includes:
[0032] Obtain the actual joint state, the actual joint torque, the quadrature-axis actual current, and the direct-axis actual current feedback by the joint motor.
[0033] According to a second aspect of the embodiments of the present disclosure, a joint motor control device is provided, and the device includes:
[0034] A current module for determining a quadrature-axis control current according to a desired joint torque;
[0035] A voltage module for inputting the quadrature-axis control current, the direct-axis control current, the quadrature-axis actual current, and the direct-axis actual current feedback by the joint motor into a reinforcement learning controller to obtain a quadrature-axis control voltage and a direct-axis control voltage output by the reinforcement learning controller;
[0036] A control module for controlling the motion of the joint motor according to the quadrature-axis control voltage and the direct-axis control voltage;
[0037] Wherein, the reinforcement learning controller is configured to execute in each control frame:
[0038] Determine a reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the quadrature-axis actual current and the direct-axis actual current of the current control frame, and add an experience point of the previous control frame to the experience pool;
[0039] Update the network parameters of the reinforcement learning model according to a parameter set of at least one experience point in the experience pool, the parameter set including the state of the current control frame, the quadrature-axis control voltage, the direct-axis control voltage, the reward value, and the state of the next control frame, and the state including the quadrature-axis actual current and the direct-axis actual current;
[0040] Generate the quadrature-axis control voltage and the direct-axis control voltage of the current control frame by using the reinforcement learning model.
[0041] In one embodiment of the present disclosure, the current module is configured to:
[0042] Determine the quadrature-axis control current according to the desired joint torque, as well as the amplitude of the permanent magnet flux linkage and the number of pole pairs of the joint motor.
[0043] In one embodiment of the present disclosure, the state further includes a quadrature-axis current error, a direct-axis current error, a quadrature-axis cumulative error, and a direct-axis cumulative error;
[0044] Wherein, the quadrature-axis current error includes the error between the quadrature-axis control current of the previous control frame and the quadrature-axis actual current of the current control frame;
[0045] The direct-axis current error includes the error between the direct-axis control current of the previous control frame and the direct-axis actual current of the current control frame;
[0046] The quadrature-axis cumulative error includes the sum of the quadrature-axis current errors of multiple control frames including the current control frame;
[0047] The direct-axis cumulative error includes the sum of the direct-axis current errors of multiple control frames including the current control frame.
[0048] In one embodiment of the present disclosure, when the reinforcement learning controller is configured to determine the reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, as well as the quadrature-axis actual current and the direct-axis actual current of the current control frame, it is configured to:
[0049] Determine the reward value according to the quadrature-axis current error and the direct-axis current error of the current control frame, as well as the quadrature-axis control voltage and the direct-axis control voltage of the previous frame.
[0050] In one embodiment of the present disclosure, the reinforcement learning controller is further configured to, at each control frame:
[0051] Smooth the network parameters of the reinforcement learning model according to the network parameters before and after the update of the reinforcement learning model.
[0052] In one embodiment of the present disclosure, when the reinforcement learning controller is configured to generate the quadrature-axis control voltage and the direct-axis control voltage by using the reinforcement learning model at each control frame, it is configured to perform any one of the following:
[0053] Before the network parameters of the reinforcement learning model are updated, generate the quadrature-axis control voltage and the direct-axis control voltage by using the reinforcement learning model;
[0054] After the network parameters of the reinforcement learning model are updated, the reinforcement learning model is used to generate a quadrature-axis control voltage and a direct-axis control voltage;
[0055] After the network parameters of the reinforcement learning model are smoothed, the reinforcement learning model is used to generate a quadrature-axis control voltage and a direct-axis control voltage.
[0056] In an embodiment of the present disclosure, the device further includes a torque module, and the torque module is configured to:
[0057] Input the desired joint state in the motion control instruction and the actual joint state feedback by the joint motor into an impedance controller to obtain the desired joint torque output by the impedance controller.
[0058] In an embodiment of the present disclosure, the desired joint state includes a desired joint position and a desired joint speed; and / or,
[0059] The actual joint state includes an actual joint position and an actual joint speed.
[0060] In an embodiment of the present disclosure, the device further includes an acquisition module, and the acquisition module is configured to:
[0061] Acquire the actual joint state, the actual joint torque, the quadrature-axis actual current, and the direct-axis actual current feedback by the joint motor.
[0062] According to a third aspect of the embodiments of the present disclosure, there is provided a robot, which includes a memory and a processor. The memory is used to store computer instructions that can run on the processor, and the processor is configured to implement the joint motor control method described in the first aspect when executing the computer instructions.
[0063] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.
[0064] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0065] The joint motor control method provided by the embodiments of the present disclosure first determines the quadrature-axis control current according to the desired joint torque; then inputs the quadrature-axis control current, the direct-axis control current, the actual quadrature-axis current and the actual direct-axis current feedback by the joint motor into a reinforcement learning controller to obtain the quadrature-axis control voltage and the direct-axis control voltage output by the reinforcement learning controller; finally, controls the joint motor to move according to the quadrature-axis control voltage and the direct-axis control voltage. Since the reinforcement learning controller is used to execute in each control frame: determine a reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the actual quadrature-axis current and the actual direct-axis current of the current control frame, and add the experience points of the previous control frame to the experience pool; update the network parameters of the reinforcement learning model according to the parameter sets of at least one experience point in the experience pool, the parameter sets include the state of the current control frame, the quadrature-axis control voltage, the direct-axis control voltage, the reward value and the state of the next control frame, and the state includes the actual quadrature-axis current and the actual direct-axis current; use the reinforcement learning model to generate the quadrature-axis control voltage and the direct-axis control voltage of the current control frame, so the accuracy of the joint motor in the control process can be improved, and the control precision and motion complexity of the robot can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0067] Figure 1 is a flowchart of a joint motor control method shown in an exemplary embodiment of the present disclosure;
[0068] Figure 2 is a schematic diagram of the operation logic of a reinforcement learning controller shown in an exemplary embodiment of the present disclosure;
[0069] Figure 3 is a flowchart of a joint motor control method shown in an exemplary embodiment of the present disclosure;
[0070] Figure 4 is a schematic structural diagram of a joint motor control device shown in an exemplary embodiment of the present disclosure;
[0071] Figure 5 is a block diagram of the structure of a robot shown in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0073] The terms used in the present disclosure are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0074] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0075] In recent years, the technology of robotics has been continuously developing, becoming more and more intelligent and automated, and the richness, stability, and flexibility of movements have all been improved to varying degrees. A bionic robot has multiple joints, and each joint is provided with a joint motor. The joint motor can drive the parts on both sides of the joint to perform relative movements. It is precisely by the movements of these joint motors that the robot can complete various actions. However, in related technologies, the accuracy of controlling the joint motors of the robot is relatively poor, resulting in the robot being unable to complete relatively complex actions.
[0076] Based on this, in a first aspect, at least one embodiment of the present disclosure provides a method for controlling a joint motor. Please refer to the appended Figure 1 , which shows the flow of the method, including steps S101 to step S103.
[0077] Among them, this method can be applied to a robot, such as a legged robot like a bipedal robot or a quadruped robot; the robot has multiple joints, and each joint is provided with a joint motor. This method can be applied to each joint motor of the robot, that is, using this method to drive the joint motor to move to complete the desired joint state obtained by the upper-level motion control according to the desired action.
[0078] In step S101, the quadrature-axis control current is determined according to the desired joint torque.
[0079] Among them, the desired joint torque can be determined in the following manner in advance: input the desired joint state in the motion control command and the actual joint state feedback by the joint motor into the impedance controller to obtain the desired joint torque output by the impedance controller.
[0080] The desired joint state can be the desired state of the joint controlled by this method determined by the upper-level motion control according to the desired motion of the robot. The joint motor can feedback the joint state at a certain frequency during motion, and this state is the actual joint state. Therefore, the actual joint state feedback by the joint motor can be obtained before executing this step. Exemplarily, the desired joint state includes the desired joint position and the desired joint velocity, and the actual joint state includes the actual joint position and the actual joint velocity. The position error can be determined according to the desired joint position and the actual joint position, and the velocity error can be determined according to the desired joint velocity and the actual joint velocity. Then, the position error and the velocity error are input into the impedance controller; the impedance controller can determine the desired joint torque based on the position error and the velocity error, combined with its internal parameters and the feedforward torque, etc.
[0081] Exemplarily, in this step, the quadrature-axis control current can be determined according to the desired joint torque, the amplitude of the permanent magnet flux linkage of the joint motor, and the number of pole pairs. For example, the quadrature-axis control current is determined according to Equation 1 below:
[0082]
[0083] In the above formula, is the desired joint torque, is the quadrature-axis control current, ψ f is the amplitude of the permanent magnet flux linkage, n p is the number of pole pairs.
[0084] In step S102, the quadrature-axis control current, the direct-axis control current, the quadrature-axis actual current feedback by the joint motor, and the direct-axis actual current are input into the reinforcement learning controller to obtain the quadrature-axis control voltage and the direct-axis control voltage output by the reinforcement learning controller.
[0085] The reinforcement learning controller is used to execute in each control frame:
[0086] Determine the reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the quadrature-axis actual current and the direct-axis actual current of the current control frame, and add the experience point of the previous control frame to the experience pool;
[0087] Update the network parameters of the reinforcement learning model according to the parameter set of at least one experience point in the experience pool, where the parameter set includes the state of the current control frame, the quadrature-axis control voltage, the direct-axis control voltage, the reward value, and the state of the next control frame;
[0088] Use the reinforcement learning model to generate the quadrature-axis control voltage and the direct-axis control voltage of the current control frame.
[0089] Exemplarily, the state includes the quadrature-axis actual current, the direct-axis actual current, the quadrature-axis current error, the direct-axis current error, the quadrature-axis cumulative error, and the direct-axis cumulative error; the quadrature-axis current error includes the error between the quadrature-axis control current of the previous control frame and the quadrature-axis actual current of the current control frame (this quadrature-axis current error is hereinafter referred to as the quadrature-axis current error of the current control frame); the direct-axis current error includes the error between the direct-axis control current of the previous control frame and the direct-axis actual current of the current control frame (this direct-axis current error is hereinafter referred to as the direct-axis current error of the current control frame); the quadrature-axis cumulative error includes the sum of the quadrature-axis current errors of multiple control frames including the current control frame (this quadrature-axis current error is hereinafter referred to as the quadrature-axis cumulative error of the current control frame); the direct-axis cumulative error includes the sum of the direct-axis current errors of multiple control frames including the current control frame (this quadrature-axis current error is hereinafter referred to as the direct-axis cumulative error of the current control frame). That is, the state of control frame t is shown in Equation 2 below:
[0090] {i d,q ,e d,q ,∫e d,q} t
[0091] In the above formula, i d,q is the direct-axis actual current and the quadrature-axis actual current, e d,q is the direct-axis current error and the quadrature-axis current error, ∫e d,q is the direct-axis cumulative error and the quadrature-axis cumulative error.
[0092] Exemplarily, the reinforcement learning controller can be used to determine the reward value of the previous control frame according to the quadrature-axis current error and the direct-axis current error of the current control frame, and the quadrature-axis control voltage and the direct-axis control voltage of the previous frame. For example, the reward value is determined according to the reward function shown in Equation 3 below:
[0093]
[0094] In the above formula, r t-1 is the reward value of the previous control frame, e d is the direct-axis current error of the current control frame, e q is the quadrature-axis current error of the current control frame, u t-1 is the quadrature-axis control voltage and the direct-axis control voltage of the previous control frame.
[0095] Exemplarily, the state of the previous control frame (i.e., the quadrature-axis actual current, direct-axis actual current, quadrature-axis current error, direct-axis current error, direct-axis cumulative error, and quadrature-axis cumulative error of the previous control frame), the quadrature-axis control voltage, the direct-axis control voltage, the reward value, and the state of the current control frame (i.e., the quadrature-axis actual current, direct-axis actual current, quadrature-axis current error, direct-axis current error, direct-axis cumulative error, and quadrature-axis cumulative error of the current control frame) are used as the experience points of the previous control frame.
[0096] Exemplarily, the network parameters of the Critic network in the reinforcement learning model can be updated in the following manner:
[0097] First, the target network value of the Critic network is determined according to Equation 4 below:
[0098] y t = r t ({i d,q , e d,q , ∫e d,q} t , U d,q ) + γQ'({i d,q , e d,q , ∫e d,q} t+1 , μ({i d,q , e d,q , ∫e d,q} t+1 |θ μ )|θ Q )
[0099] In the above formula, 0 < γ ≤ 1 is the discount factor, y t is the target network value of the Critic network at control frame t, r t is the reward value at control frame t, Q' is the Critic network before update, μ is the Actor network before update, θ Q is the network parameter of the Critic network before update, θ μ is the network parameter of the Actor network before update.
[0100] Next, the network parameters of the Critic network are updated by minimizing the loss function L shown in Equation 5 below:
[0101]
[0102] In the above formula, Q is the updated Critic network, which is used to calculate the evaluation value of the state-action pair, θ Qis the network parameter of the updated Critic network, and M is the number of the at least one experience point in this step.
[0103] That is to say, in this method, the target network value of the Critic network at each experience point in the at least one experience point is determined according to the above formula 4, and then according to the above formula 5, the target network value of the Critic network at each experience point in the at least one experience point is used to update the network parameters of the Critic network.
[0104] As another example, the network parameters of the Actor network in the reinforcement learning model can be updated according to the gradient ascent method shown in the following formula 6:
[0105]
[0106] In the above formula, J(θ μ )=E[Q({i d,q , e d,q , ∫e d,q} t , U d,q |U d,q ]=μ({i d,q , e d,q , ∫e d,q} t )
[0107] It should be understood that the network parameters of the reinforcement learning model can also be smoothed according to the network parameters before and after the reinforcement learning model is updated. For example, the network parameters of the Critic network and the network parameters of the Actor network in the reinforcement learning model are smoothed in the manner shown in the following formula 7:
[0108]
[0109] In the above formula, θ Q″ is the network parameter after smoothing of the Critic network, θ Q is the network parameter before the Critic network update, θ Q′ is the updated network parameter of the Critic network, θ μ″ is the network parameter after smoothing of the Actor network, θ μ is the network parameter before the Actor network update, θ μ′ Updated network parameters for the Actor network.
[0110] Based on this, the reinforcement learning controller is used for any of the following when generating the quadrature-axis control voltage and the direct-axis control voltage using the reinforcement learning model in each control frame:
[0111] Before updating the network parameters of the reinforcement learning model, use the reinforcement learning model to generate a quadrature-axis control voltage and a direct-axis control voltage;
[0112] After updating the network parameters of the reinforcement learning model, use the reinforcement learning model to generate a quadrature-axis control voltage and a direct-axis control voltage;
[0113] After smoothing the network parameters of the reinforcement learning model, use the reinforcement learning model to generate a quadrature-axis control voltage and a direct-axis control voltage.
[0114] For example, the above three methods can all generate the quadrature-axis control voltage and the direct-axis control voltage of the current control frame in the manner shown in Equation 7 below:
[0115] U d,q =μ({i d,q ,e d,q ,∫e d,q} t |θ μ )+ε t
[0116] In the above formula, ε t is a noise network, μ is an Actor network, and θ μ (the network parameters of the Actor network) are different in the above three methods.
[0117] It can be understood that the above operations performed by the reinforcement learning controller in each control frame can be regarded as a cycle in the training process of the reinforcement learning model. Refer to the appendix Figure 2 , which shows the operation logic of this reinforcement learning controller. The direct and quadrature-axis control currents and the direct and quadrature-axis control currents i d、q are input into the controller in the form of errors for signal processing, so as to obtain the empirical points of the previous control frame. Then, the reinforcement learning agent (i.e., the reinforcement learning network) updates the network parameters and determines the actions according to the parameter set of at least one empirical point (i.e., determines the quadrature-axis control voltage and the direct-axis control voltage of the current control frame). Finally, after the controller applies the actions to the environment, a feedback value is generated (i.e., after applying the quadrature-axis control voltage and the direct-axis control voltage to the control of the joint motor, the direct and quadrature-axis control currents i d、q ) are generated.
[0118] In step S103, control the joint motor to move according to the quadrature-axis control voltage and the direct-axis control voltage.
[0119] The joint motor control method provided by the embodiments of the present disclosure first determines the quadrature-axis control current according to the desired joint torque; then inputs the quadrature-axis control current, the direct-axis control current, the quadrature-axis actual current and the direct-axis actual current fed back by the joint motor into a reinforcement learning controller to obtain the quadrature-axis control voltage and the direct-axis control voltage output by the reinforcement learning controller; finally, controls the joint motor to move according to the quadrature-axis control voltage and the direct-axis control voltage. Since the reinforcement learning controller is used to execute in each control frame: determine a reward value according to the quadrature-axis control current and the direct-axis control current in the previous control frame, and the quadrature-axis actual current and the direct-axis actual current in the current control frame, and add an experience point of the previous control frame to the experience pool; update the network parameters of the reinforcement learning model according to the parameter sets of at least one experience point in the experience pool, the parameter sets include the state in the current control frame, the quadrature-axis control voltage, the direct-axis control voltage, the reward value and the state in the next control frame, and the state includes the quadrature-axis actual current and the direct-axis actual current; generate the quadrature-axis control voltage and the direct-axis control voltage in the current control frame by using the reinforcement learning model, so the accuracy of the joint motor in the control process can be improved, and the control precision and motion complexity of the robot can be improved.
[0120] Please refer to the attached Figure 3 , which exemplarily shows the flowchart of the joint motor control method obtained by combining the above-mentioned multiple embodiments. Among them, the upper-level motion control issues the desired joint position θ d , the desired joint speed v d , the impedance controller parameters K p , K d and the feedback torque τ ff . Combining the actual joint position θ and the actual joint speed v fed back by the joint motor, the desired joint torque is obtained through the impedance controller The quadrature-axis control current is obtained through the relationship between torque and current (such as the relationship shown in Equation 1 above) Combining with the preset direct-axis control current (such as ), combining the quadrature-axis actual current i q and the direct-axis actual current i d fed back by the joint motor, the quadrature-axis and direct-axis control voltages are obtained through the reinforcement learning controller The three-phase voltage is obtained through pulse width modulation and the motor is driven to operate through an inverter.
[0121] In this method, the control of the joint motor by the controller can be regarded as the training process of the reinforcement learning network. This process can adapt to various working conditions and environments, so as to improve the robustness and adaptability of the joint motor. Moreover, the controller is simply designed and has strong transplantability.
[0122] According to a second aspect of the embodiments of the present disclosure, a joint motor control device is provided. Please refer to the appended Figure 4 , the device includes:
[0123] A current module 401, configured to determine a quadrature-axis control current according to a desired joint torque;
[0124] A voltage module 402, configured to input the quadrature-axis control current, the direct-axis control current, the quadrature-axis actual current and the direct-axis actual current fed back by the joint motor into a reinforcement learning controller, and obtain the quadrature-axis control voltage and the direct-axis control voltage output by the reinforcement learning controller;
[0125] A control module 403, configured to control the joint motor to move according to the quadrature-axis control voltage and the direct-axis control voltage;
[0126] Wherein, the reinforcement learning controller is configured to execute at each control frame:
[0127] Determine a reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the quadrature-axis actual current and the direct-axis actual current of the current control frame, and add an experience point of the previous control frame to the experience pool;
[0128] Update the network parameters of the reinforcement learning model according to a parameter set of at least one experience point in the experience pool, the parameter set includes the state of the current control frame, the quadrature-axis control voltage, the direct-axis control voltage, the reward value and the state of the next control frame, and the state includes the quadrature-axis actual current and the direct-axis actual current;
[0129] Generate the quadrature-axis control voltage and the direct-axis control voltage of the current control frame by using the reinforcement learning model.
[0130] In an embodiment of the present disclosure, the current module is configured to:
[0131] Determine a quadrature-axis control current according to a desired joint torque, and the amplitude of the permanent magnet flux linkage and the number of pole pairs of the joint motor.
[0132] In an embodiment of the present disclosure, the state further includes a quadrature-axis current error, a direct-axis current error, a quadrature-axis cumulative error and a direct-axis cumulative error;
[0133] Wherein, the quadrature-axis current error includes the error between the quadrature-axis control current of the previous control frame and the quadrature-axis actual current of the current control frame;
[0134] The direct-axis current error includes the error between the direct-axis control current of the previous control frame and the direct-axis actual current of the current control frame;
[0135] The quadrature-axis cumulative error includes the sum of the quadrature-axis current errors of multiple control frames including the current control frame;
[0136] The direct-axis cumulative error includes the sum of the direct-axis current errors of multiple control frames including the current control frame.
[0137] In an embodiment of the present disclosure, when the reinforcement learning controller is used to determine a reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the quadrature-axis actual current and the direct-axis actual current of the current control frame, it is used for:
[0138] Determine the reward value according to the quadrature-axis current error and the direct-axis current error of the current control frame, and the quadrature-axis control voltage and the direct-axis control voltage of the previous frame.
[0139] In an embodiment of the present disclosure, the reinforcement learning controller is further used in each control frame:
[0140] Smooth the network parameters of the reinforcement learning model according to the network parameters before and after the update of the reinforcement learning model.
[0141] In an embodiment of the present disclosure, when the reinforcement learning controller is used to generate a quadrature-axis control voltage and a direct-axis control voltage using the reinforcement learning model in each control frame, it is used for any one of the following:
[0142] Before the network parameters of the reinforcement learning model are updated, use the reinforcement learning model to generate a quadrature-axis control voltage and a direct-axis control voltage;
[0143] After the network parameters of the reinforcement learning model are updated, use the reinforcement learning model to generate a quadrature-axis control voltage and a direct-axis control voltage;
[0144] After the network parameters of the reinforcement learning model are smoothed, use the reinforcement learning model to generate a quadrature-axis control voltage and a direct-axis control voltage.
[0145] In an embodiment of the present disclosure, the device further includes a torque module, and the torque module is used for:
[0146] Input the desired joint state in the motion control command and the actual joint state fed back by the joint motor into the impedance controller to obtain the desired joint torque output by the impedance controller.
[0147] In an embodiment of the present disclosure, the desired joint state includes a desired joint position and a desired joint speed; and / or,
[0148] The actual joint state includes an actual joint position and an actual joint speed.
[0149] In an embodiment of the present disclosure, the device further includes an acquisition module, and the acquisition module is used for:
[0150] Obtain the actual joint state, the actual joint torque, the quadrature-axis actual current, and the direct-axis actual current feedback from the joint motor.
[0151] In a third aspect, at least one embodiment of the present disclosure provides a robot. Please refer to the appended Figure 5 figure, which shows the structure of the robot. The robot includes a memory and a processor. The memory is used to store computer instructions that can be run on the processor, and the processor is used to control the joint motor based on the method according to any one of the first aspect when executing the computer instructions.
[0152] In a fourth aspect, at least one embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method according to any one of the first aspect is implemented.
[0153] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0154] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for controlling a joint motor, characterized in that, The method includes: Determining the quadrature-axis control current according to the desired joint torque; Inputting the quadrature-axis control current, the direct-axis control current, the actual quadrature-axis current and the actual direct-axis current fed back by the joint motor into a reinforcement learning controller to obtain the quadrature-axis control voltage and the direct-axis control voltage output by the reinforcement learning controller; Controlling the motion of the joint motor according to the quadrature-axis control voltage and the direct-axis control voltage; Wherein, the reinforcement learning controller is used to execute in each control frame: Determining a reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the actual quadrature-axis current and the actual direct-axis current of the current control frame, and adding the experience points of the previous control frame to the experience pool; Updating the network parameters of the reinforcement learning model according to the parameter sets of at least one experience point in the experience pool, the parameter sets including the state of the current control frame, the quadrature-axis control voltage, the direct-axis control voltage, the reward value and the state of the next control frame, and the state including the actual quadrature-axis current and the actual direct-axis current; Generating the quadrature-axis control voltage and the direct-axis control voltage of the current control frame by using the reinforcement learning model.
2. The joint motor control method according to claim 1, wherein The determining the quadrature-axis control current according to the desired joint torque includes: Determining the quadrature-axis control current according to the desired joint torque, the amplitude of the permanent magnet flux linkage of the joint motor and the number of pole pairs.
3. The joint motor control method according to claim 1, wherein The state further includes the quadrature-axis current error, the direct-axis current error, the quadrature-axis cumulative error and the direct-axis cumulative error; Wherein, the quadrature-axis current error includes the error between the quadrature-axis control current of the previous control frame and the actual quadrature-axis current of the current control frame; The direct-axis current error includes the error between the direct-axis control current of the previous control frame and the actual direct-axis current of the current control frame; The quadrature-axis cumulative error includes the sum of the quadrature-axis current errors of multiple control frames including the current control frame; The direct-axis cumulative error includes the sum of the direct-axis current errors of multiple control frames including the current control frame.
4. The joint motor control method according to claim 3, wherein When the reinforcement learning controller is used to determine the reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the actual quadrature-axis current and the actual direct-axis current of the current control frame, it is used to: Determining the reward value according to the quadrature-axis current error and the direct-axis current error of the current control frame, and the quadrature-axis control voltage and the direct-axis control voltage of the previous frame.
5. The joint motor control method according to claim 1, wherein The reinforcement learning controller is further used to in each control frame: Smoothing the network parameters of the reinforcement learning model according to the network parameters before and after the update of the reinforcement learning model.
6. The joint motor control method according to claim 5, wherein When the reinforcement learning controller is used to generate the quadrature-axis control voltage and the direct-axis control voltage by using the reinforcement learning model in each control frame, it is used for any one of the following: Before the network parameters of the reinforcement learning model are updated, generating the quadrature-axis control voltage and the direct-axis control voltage by using the reinforcement learning model; After the network parameters of the reinforcement learning model are updated, generating the quadrature-axis control voltage and the direct-axis control voltage by using the reinforcement learning model; After the network parameters of the reinforcement learning model are smoothed, generating the quadrature-axis control voltage and the direct-axis control voltage by using the reinforcement learning model.
7. The joint motor control method according to claim 1, characterized in that The method further includes: Input the desired joint state in the motion control instruction and the actual joint state feedback from the joint motor into an impedance controller to obtain the desired joint torque output by the impedance controller.
8. The joint motor control method according to claim 7, wherein The desired joint state includes a desired joint position and a desired joint velocity; and / or, The actual joint state includes an actual joint position and an actual joint velocity.
9. The joint motor control method according to claim 7, wherein The method further includes: Obtain the actual joint state, the actual joint torque, the actual quadrature-axis current, and the actual direct-axis current feedback from the joint motor.
10. An articulated motor control device, characterized in that, The device includes: A current module for determining a quadrature-axis control current according to the desired joint torque; A voltage module for inputting the quadrature-axis control current, the direct-axis control current, the actual quadrature-axis current, and the actual direct-axis current feedback from the joint motor into a reinforcement learning controller to obtain the quadrature-axis control voltage and the direct-axis control voltage output by the reinforcement learning controller; A control module for controlling the motion of the joint motor according to the quadrature-axis control voltage and the direct-axis control voltage; Wherein, the reinforcement learning controller is configured to execute in each control frame: Determine a reward value according to the quadrature-axis control current and the direct-axis control current of the previous control frame, and the actual quadrature-axis current and the actual direct-axis current of the current control frame, and add an experience point of the previous control frame to the experience pool; Update the network parameters of the reinforcement learning model according to the parameter sets of at least one experience point in the experience pool, the parameter sets including the state of the current control frame, the quadrature-axis control voltage, the direct-axis control voltage, the reward value, and the state of the next control frame, the state including the actual quadrature-axis current and the actual direct-axis current; Generate the quadrature-axis control voltage and the direct-axis control voltage of the current control frame by using the reinforcement learning model.
11. A robot, characterized in that, The robot includes a memory and a processor. The memory is used for storing computer instructions that can be run on the processor, and the processor is used for implementing the joint motor control method according to any one of claims 1 to 9 when executing the computer instructions.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 9.