Dual-motor seamless variable stiffness control method based on deep reinforcement learning and related equipment
By modeling the dual-motor variable stiffness control task as a Markov decision process and using deep reinforcement learning to optimize the policy network, an intelligent control network is generated. This solves the dependence of variable stiffness control methods on precise mathematical models and achieves optimal control in unknown or dynamic environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIAXING MINSHUO INTELLIGENT TECH CO LTD
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-17
AI Technical Summary
Existing variable stiffness control methods rely on precise mathematical models, making it difficult to achieve optimal performance in unknown or dynamically changing environments, and require complex online optimization or parameter tuning.
A seamless variable stiffness control method for dual motors based on deep reinforcement learning is adopted. The dual motor variable stiffness control task is modeled as a Markov decision process. The policy network is optimized at the proximal end through the state space, action space and variable stiffness reward function to generate an intelligent control network and realize the autonomous learning of motor torque commands.
It eliminates the reliance on precise mathematical models and can autonomously learn optimal control strategies in unknown or dynamic environments, achieving a balance in variable stiffness control behavior and avoiding complex stiffness mapping or distributive law design.
Smart Images

Figure CN121887010A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent control technology, and in particular to a seamless variable stiffness control method and related equipment for dual motors based on deep reinforcement learning. Background Technology
[0002] Currently, variable stiffness control relies on a precise mathematical model for multi-motor coordinated control. The accuracy of the model and the setting of parameters will affect the effect of stiffness control. Some control methods rely on rapid feedback from sensors, and online adjustment of variable stiffness control is time-consuming and unstable.
[0003] In real-world systems, strong nonlinear factors such as gear backlash, friction, and time-varying parameters make accurate modeling extremely difficult, leading to a decline in the performance of model-based control methods.
[0004] Furthermore, these methods typically require complex online optimization or parameter tuning, making it difficult to achieve optimal performance in unknown or dynamically changing environments. Summary of the Invention
[0005] The main objective of this application is to propose a seamless variable stiffness control method and related equipment for dual motors based on deep reinforcement learning. This method aims to eliminate the reliance on precise mathematical models, eliminate the need for manually designing complex stiffness mappings or allocation laws, and enable the system to proactively learn the optimal control strategy through interaction with the environment to achieve variable stiffness control behavior.
[0006] To achieve the above objectives, one aspect of this application proposes a seamless variable stiffness control method for dual motors based on deep reinforcement learning. The dual motors include a first motor and a second motor, which are coupled together to output to a single output shaft. The method includes the following steps: The dual-motor variable stiffness control task is modeled as a Markov decision process, resulting in a state space, an action space, and a variable stiffness reward function. The state space is a set of the desired state of the task, the interaction state with the environment, and the motion states of the dual motors and the output shaft. By performing proximal policy optimization on the policy network using the state space, the action space, and the variable stiffness reward function, the policy network learns to generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thus obtaining an intelligent control network. In each control cycle, the current state vector of the dual motors is collected, and the current state vector is input into the intelligent control network for control decision-making to obtain the optimal motor torque command; The dual motors are driven to perform variable stiffness coordinated motion according to the optimal motor torque command.
[0007] In some embodiments, the state space is tuple data, the task expectation state includes expectation stiffness, expectation angle and expectation velocity, the environmental interaction state is external interaction force, and the motion state includes the first angle and first velocity of the first motor, the second angle and second velocity of the second motor, and the third angle and third velocity of the output shaft. The motion space includes the first torque increment of the first motor and the second torque increment of the second motor; The variable stiffness reward function is a multi-objective composite reward function, including a trajectory tracking term, a stiffness realization term, an energy benefit term, and a smoothing performance term. The trajectory tracking term is used to penalize trajectory tracking errors, the stiffness realization term is used to control the relationship between force and displacement, the energy benefit term is used to minimize energy consumption, and the smoothing performance term is used to penalize incremental abrupt changes.
[0008] In some embodiments, obtaining the variable stiffness reward function includes the following steps: The trajectory error is quantized based on the first difference between the third angle and the desired angle and the second difference between the third velocity and the desired velocity to obtain the trajectory tracking term; A virtual spring simulation is performed based on the first difference and the desired stiffness to obtain a stiffness model. The spring simulation error is then quantified based on the external interaction force and the stiffness model to obtain the stiffness realization term. The energy efficiency term is obtained by performing comprehensive amplitude quantization based on the first torque increment and the second torque increment. Torque increment quantization is performed based on the first torque increment and the second torque increment to obtain the smoothing performance term; The variable stiffness reward function is obtained by weighting the trajectory tracking term, the stiffness realization term, the energy benefit term, and the smoothness performance term.
[0009] In some embodiments, the expression for the variable stiffness reward function is: ; in, α , β , c , d , or These are the weighting coefficients for each item. r t As a reward value, i 1 represents the first angle. For the first speed, i 2 represents the second angle. The second speed, i l For the third angle, The third velocity, F ext For the external interaction force, K d For the desired stiffness, i ref For the desired angle, For the desired speed, △T 1 represents the first torque increment. △T 2 represents the second torque increment.
[0010] In some embodiments, the process of performing proximal policy optimization on the policy network using the state space, the action space, and the variable stiffness reward function, enabling the policy network to learn and generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thereby obtaining an intelligent control network, includes the following steps: Initialize the policy network and interact with the environment through the initialized policy network to obtain training data, wherein the training data includes state data, action data and reward data; Based on the training data, a generalized dominance estimate is performed on the policy network to obtain the dominance value; The training data is randomly sampled based on a preset first quantity to obtain sample data; The policy loss function is determined based on the advantage value, and the policy network is optimized by limiting the update magnitude using the policy loss function based on the sample data to obtain the optimized policy network. Then, the process returns to the step of randomly sampling the training data based on a preset first number to obtain sample data, until the training termination condition is met. Then, the last obtained policy network is confirmed as the intelligent control network.
[0011] In some embodiments, the step of acquiring the current state vector of the dual motors in each control cycle, inputting the current state vector into the strategy network for control decision-making, and obtaining the optimal motor torque command includes the following steps: In each control cycle, the current state vectors of the dual motors and the output shaft are acquired according to the state space; The current state vector is input into the policy network for forward propagation to obtain the optimal action, wherein the optimal action includes a first optimal torque increment action and a second optimal torque increment action. The optimal motor torque command from the previous control cycle is updated based on the optimal action to obtain the updated optimal motor torque command.
[0012] In some embodiments, the optimal motor torque command includes a first motor torque command and a second motor torque command. The step of updating the optimal motor torque command of the previous control cycle based on the optimal action to obtain the updated optimal motor torque command includes the following steps: The first motor torque command of the previous control cycle is updated according to the first optimal torque increment action to obtain the updated first motor torque command. The second motor torque command of the previous control cycle is updated according to the second optimal torque increment action to obtain the updated second motor torque command.
[0013] To achieve the above objectives, another aspect of this application proposes a seamless variable stiffness control system for dual motors based on deep reinforcement learning. The system is used to implement the above method and includes: The Markov decision module is used to model the dual-motor variable stiffness control task as a Markov decision process to obtain the state space, action space and variable stiffness reward function. The state space is a set of the task expectation state, the environmental interaction state and the motion state of the dual motors and the output shaft. The strategy network training module is used to perform proximal policy optimization on the strategy network through the state space, the action space and the variable stiffness reward function, so that the strategy network learns to generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thereby obtaining an intelligent control network. The control decision module is used to collect the current state vector of the dual motors in each control cycle, input the current state vector into the intelligent control network for control decision-making, and obtain the optimal motor torque command. The motor drive module is used to drive the two motors to perform variable stiffness coordinated motion according to the optimal motor torque command.
[0014] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0015] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0016] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0017] The embodiments of this application include at least the following beneficial effects: This application provides a seamless variable stiffness control method and related equipment for dual motors based on deep reinforcement learning. This scheme models the dual-motor variable stiffness control task as a Markov decision process, obtaining a state space, action space, and variable stiffness reward function. Proximal policy optimization is performed on the policy network using the state space, action space, and variable stiffness reward function, enabling the policy network to learn and generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thus obtaining an intelligent control network. In each control cycle, the current state vectors of the dual motors are collected and input into the intelligent control network for control decision-making, obtaining the optimal motor torque command. The dual motors are then driven to perform variable stiffness cooperative motion according to the optimal motor torque command. The embodiments of this application can eliminate the dependence on precise mathematical models, eliminating the need for manual design of stiffness mappings or distributive laws, and actively learn the optimal control strategy through interaction with the environment to achieve variable stiffness control behavior. Attached Figure Description
[0018] Figure 1 This is a flowchart of the seamless variable stiffness control method for dual motors based on deep reinforcement learning provided in the embodiments of this application; Figure 2 This is a flowchart of a seamless variable stiffness control method for dual motors based on deep reinforcement learning, provided in another embodiment of this application; Figure 3 This is a flowchart of the intelligent control network training provided in the embodiments of this application; Figure 4 This is a flowchart of the deep reinforcement learning and control method for a dual-motor system provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a seamless variable stiffness control system for dual motors based on deep reinforcement learning provided in an embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] This application provides a method and related equipment for seamless variable stiffness control of dual motors based on deep reinforcement learning. The scheme models the dual-motor variable stiffness control task as a Markov decision process, obtaining a state space, action space, and variable stiffness reward function. Proximal policy optimization is performed on the policy network using the state space, action space, and variable stiffness reward function, enabling the policy network to learn and generate motor torque commands that achieve variable stiffness behavior balance according to the environment, resulting in an intelligent control network. In each control cycle, the current state vectors of the dual motors are collected and input into the intelligent control network for control decision-making, obtaining the optimal motor torque command. The dual motors are then driven to perform variable stiffness cooperative motion according to the optimal motor torque command. This application embodiment can overcome the dependence on precise mathematical models and actively learn the optimal control strategy through interaction with the environment to achieve variable stiffness control behavior.
[0022] The seamless variable stiffness control method for dual motors based on deep reinforcement learning provided in this application relates to the field of intelligent control. This method can be applied to terminals, servers, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the seamless variable stiffness control method for dual motors based on deep reinforcement learning, but is not limited to the above forms.
[0023] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0024] Figure 1 This is an optional flowchart of the seamless variable stiffness control method for dual motors based on deep reinforcement learning provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S104.
[0025] Step S101: Model the dual-motor variable stiffness control task as a Markov decision process to obtain the state space, action space and variable stiffness reward function.
[0026] Step S102: The policy network is optimized for near-end policy through state space, action space and variable stiffness reward function, so that the policy network learns to generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thus obtaining the intelligent control network.
[0027] Step S103: In each control cycle, the current state vector of the two motors is collected, and the current state vector is input into the intelligent control network for control decision-making to obtain the optimal motor torque command.
[0028] Step S104: Drive the two motors to perform variable stiffness coordinated motion according to the optimal motor torque command.
[0029] In this embodiment, the animal can effectively complete complex variable stiffness movements through muscles. During the contraction and extension of biological muscles, the coordinated work of antagonistic and synergistic muscles not only changes stiffness but also overcomes conditions such as muscle gaps. For example, during forward or reverse movement of the antagonistic muscle, synergistic muscles simultaneously contribute to balance and stabilize the movement, enabling the muscle movement to achieve the variable stiffness characteristics required by the neuron. Inspired by this, this embodiment utilizes dual-motor cooperative control instead of a single motor. The dual motors include a first motor and a second motor, which are coupled together to output to a single output shaft. Through dual-motor cooperative control, the motors themselves acquire variable stiffness characteristics.
[0030] Deep reinforcement learning, as an advanced intelligent control method, can autonomously learn the optimal control strategy through interaction with the environment without the need for an explicit system model. It provides a new technical approach to solving the current challenge of variable stiffness control relying on precise mathematical models. The data-driven intelligent variable stiffness control paradigm can overcome the dependence of related technologies on precise mathematical models.
[0031] Specifically, a physical dual-motor cooperative drive system can be abstracted as a learning problem involving a continuous interaction between an agent and a dynamic environment. Therefore, the dual-motor variable stiffness control task is modeled as a Markov decision process. In reinforcement learning, the state space, action space, and reward function are its core components, which together construct the framework for the agent's interaction with the environment, driving the agent from trial-and-error learning to optimal decision-making.
[0032] The state space defines all the environmental information that the agent can perceive at each decision moment. In this embodiment, the state space is a set of the task expectation state, the environmental interaction state, and the motion states of the two motors and the output shaft.
[0033] For example, state space S For 10-tuple data, it is defined as: (1); in, t For each time step, the desired task state includes task instructions planned by the task planner, such as the desired stiffness. K d Expectation perspective i ref and expected speed The environmental interaction state is the external interaction force collected by the torque sensor. F ext Such as ground reaction force; motion state includes the first angle of the first motor. i 1 and first speed The second angle of the second motor i 2 and the second speed The third angle of the output shaft i l With the third speed .
[0034] Action space is the set of operations that an agent can perform in a specific state. In this embodiment, action space can be visualized as the control commands that the agent outputs to the two motors, and incremental control is adopted to ensure stability.
[0035] For example, the action space A is defined as: (2); in, △T 1 represents the first torque increment of the first motor. △T 2 represents the second torque increment of the second motor.
[0036] The variable stiffness reward function is a multi-objective reward function used to evaluate the quality of an action. In this embodiment, the variable stiffness reward function is used to take into account the optimal strategy of tracking, stiffness and energy efficiency. It not only guides the agent to learn how to track with high precision during the interaction process, but also achieves the best balance between tracking accuracy, stiffness compliance and energy efficiency, so that the agent can realize adjustable and variable functions such as stiffness, energy consumption and tracking according to actual needs.
[0037] For example, the variable stiffness reward function includes a trajectory tracking term, a stiffness realization term, an energy benefit term, and a smoothing performance term, which can be defined as: (3); in, oh 1. oh 2. oh 3. oh 4 represents the weighting coefficient for each item. r t As a reward value, r track For trajectory tracking items, r stifness For stiffness realization, r energy For energy efficiency items, r smooth This is a smoothing performance term.
[0038] Optionally, the variable stiffness reward function can be obtained through steps S201 to S205.
[0039] Step S201: Perform trajectory error quantization based on the first difference between the third angle and the desired angle and the second difference between the third velocity and the desired velocity to obtain the trajectory tracking term.
[0040] Step S202: Perform virtual spring simulation based on the first difference and the desired stiffness to obtain a stiffness model, and quantify the spring simulation error based on the external interaction force and the stiffness model to obtain the stiffness realization term.
[0041] Step S203: Perform comprehensive amplitude quantization based on the first torque increment and the second torque increment to obtain the energy benefit term.
[0042] Step S204: Quantize the torque increment based on the first torque increment and the second torque increment to obtain the smoothing performance term.
[0043] Step S205: The trajectory tracking term, stiffness realization term, energy benefit term, and smoothing performance term are weighted and processed to obtain the variable stiffness reward function.
[0044] Specifically, assuming all weight coefficients of the variable stiffness reward function oh 1. oh 2. oh 3. oh All four are -1. The variable stiffness reward function can also be defined as: (4); in, α , β , c , d , or These are the weighting coefficients for each item.
[0045] The following is combined with Figure 2 Explain the impact of each term in the variable stiffness reward function on the strategy.
[0046] This is a trajectory tracking term used to penalize trajectory tracking errors and encourage precise tracking. Its goal is to guide the agent's control output axis to closely follow the expected trajectory provided by the task planner; any deviation will be penalized, and this deviation will be minimized by adjusting the motor torque.
[0047] First, the angular deviation is quantified by calculating the difference between the third angle and the desired angle, resulting in the first difference. Next, the velocity deviation is quantified by calculating the difference between the third velocity and the desired velocity, resulting in the second difference. Finally, the first and second differences are combined using a weighted square method.
[0048] Weight of the first difference α Weight of the second difference β Penalize deviations by increasing their weights. α and β This allows the trajectory tracking item to be sensitive to deviations in angle and speed, ensuring tracking accuracy.
[0049] In scenarios where robots perform "pick-and-place" operations on assembly lines, increasing the weight is crucial. α and β Learn a strategy that prioritizes accuracy.
[0050] - c ( F ext - K d ( i l - i ref )) 2This is a stiffness realization term used to control the relationship between force and displacement, encouraging actual interaction forces. F ext With stiffness model K d ( i l - i ref Consistency is the core of achieving variable stiffness control. The force-displacement relationship of the entire system resembles a stiffness constant. K d The spring.
[0051] For example, when robots collaborate with humans in uncertain environments, i.e., in low-stiffness-safe interaction scenarios, then... K d The value is low, even if a large displacement is quantized based on the first difference ( i l - i ref ) Expectations are also very low. Therefore, when encountering a human body or an obstacle... F ext When the joint is enlarged, it allows for natural displacement, absorbing impact and ensuring safety.
[0052] This is an energy efficiency term used to penalize large torque outputs and minimize energy consumption, thereby improving energy efficiency by performing tasks with minimal torque. The energy efficiency term focuses on the magnitude of two torque increments. By summing the squares of the two torque increments, a penalty is imposed on the torque increment with the larger sum, encouraging the agent to explore and adopt efficient control strategies.
[0053] In scenarios where wilderness rescue robots need to perform tasks for extended periods, d The weight will be higher. The intelligent system will learn to achieve its goals using a clever dual-motor torque distribution method. For example, when the load is affected by gravity, it will learn to use gravity to assist in movement, instead of always using the motor to fight gravity.
[0054] The smoothing performance term is used to penalize sudden changes in increments to control drastic changes in commands and ensure smooth control. Although its input is the same as the energy efficiency term, its purpose is to utilize the torque increment itself to represent the change in command between the current control cycle and the previous control cycle. Therefore, the smoothing performance term essentially penalizes the magnitude of torque changes between adjacent control cycles.
[0055] For example, in a scenario where a robot is walking around carrying a glass full of water, orThe weight of the joint will be very high, and even if the torque needs to be changed, it will output a series of small and continuous torque increments instead of a huge step change. This prevents the liquid in the cup from spilling due to the sudden start or stop or shaking of the joint.
[0056] It is understandable that, in this embodiment, the weighting coefficients of each reward item... oh 1. oh 2. oh 3. oh The value of 4 is merely an example. The value of the weighting coefficient can be set according to actual needs. It can be the same or different. This application embodiment does not impose specific restrictions, as long as the best balance between tracking accuracy, stiffness compliance and energy efficiency can be achieved.
[0057] The offline training phase begins based on a predefined state space, action space, and variable stiffness reward function. Training employs the Proximal Policy Optimization (PPO) method to train a policy network. π φ ( a t | s t The training process takes place in a simulation environment or a physical system, acquiring training data through millions of interactions. In each interaction, the policy network, based on the current output action, receives a new state from the environment after executing that action, and the reward calculated by the variable stiffness reward function. The training data is used to adjust the policy network parameters. f Convergence, thus obtaining a state that can be obtained from the state s t Mapping to optimal action distribution a t The intelligent control network is trained with a strategy that yields optimal solutions for long-term cumulative rewards. This allows the network to understand the meaning of different desired stiffness commands and autonomously determine the optimal torque distribution scheme between the two motors. Consequently, under any given stiffness command, the network can coordinate the two motors in real time, ensuring that the output shaft accurately exhibits the corresponding stiffness characteristics while meeting trajectory tracking requirements, and achieving an intelligent balance among multiple objectives such as stiffness, accuracy, and energy consumption.
[0058] The trained intelligent control network is deployed to the real-time controller of the dual-motor system. In each control cycle, the actual motion state of the output shaft and the dual motors is collected in real time by sensors, and the current state vector is formed together with the expected state of the task and the interaction state of the environment according to equation (1). s t .
[0059] The current state vector is input into the intelligent control network, which makes control decisions based on the internal parameters learned during the offline training phase to achieve variable stiffness behavior balance according to the environment, and finally outputs a set of optimal motor torque commands to update the torque of the two motors.
[0060] Because the definition of the action space in this embodiment is achieved through incremental control, the optimal motor torque command output by the intelligent control network is also in incremental form. The actual control command for driving the motor needs to be superimposed with the control command of the previous cycle.
[0061] It should be noted that when driving dual motors, the two motors independently generate corresponding torques according to control commands, and output them together to a single output shaft through coupling between the two motors, realizing data-driven intelligent variable stiffness control, eliminating the dependence on precise mathematical models, and eliminating the need for manual design of complex stiffness mapping or distribution laws.
[0062] In some embodiments, step S102 may include, but is not limited to, steps S301 to S304.
[0063] Step S301: Initialize the policy network and interact with the environment through the initialized policy network to obtain training data.
[0064] Step S302: Perform generalized advantage estimation on the policy network based on the training data to obtain the advantage value.
[0065] Step S303: Randomly sample the training data based on a preset first quantity to obtain sample data.
[0066] Step S304: Determine the policy loss function based on the advantage value, and use the policy loss function to optimize the policy network by limiting the update magnitude based on the sample data to obtain the optimized policy network. Then return to the step of randomly sampling the training data based on a preset first number to obtain the sample data, until the training end condition is met, and then confirm the last obtained policy network as the intelligent control network.
[0067] In this embodiment, the detailed process of near-end strategy optimization is as follows: Figure 3 As shown, training data is first collected and the environment is interacted with for N time steps.
[0068] Specifically, policy network π φ parameters f Randomly initialized, policy network π φ Based on the status data at the current time step s t Output a Gaussian distribution and sample an action data from it. at =[ △T 1, △T 2). Environmental execution action data a t The state data is transferred to the next time step based on its inherent physical dynamics. s t+1 Calculate reward data r t The training data ( s t , a t , r t , s t+1 Store it in the experience playback buffer. Repeat the above steps until an episode ends.
[0069] Next, execute the PPO core training loop and policy network tuning. Use the current policy network. π φ Interacting with the environment, 4096 trajectory data points were collected and stored in an empirical buffer. Then, generalized advantage estimation was used to calculate the advantage value A. t The formula for calculating the dominance value is as follows: (5); in, , c =0.99, l =0.95.
[0070] Then, small batches of data are sampled from the training data stored in the experience replay buffer at a preset first number (e.g., 64) for 10 optimization iterations (Epochs).
[0071] The policy loss function is defined as follows: (6); in, It is the probability ratio of the new strategy to the old strategy. It is the estimated advantage value, and the clip function will... r t ( f ) limited to [1 Within the range of ε, 1+ε, for example, ε=0.2.
[0072] While optimizing the policy network, it is also necessary to train the value network V_ ψ ( s This makes the predictions more accurate. The value loss function uses mean squared error loss: (7); in,R t From state s t The initial actual cumulative return.
[0073] Furthermore, to encourage policy exploration and prevent the policy network from prematurely converging to a local optimum and stopping, the policy's entropy term can be added to the total loss. S And multiply by a small coefficient of 0.01.
[0074] Entropy in policy optimization encourages exploration by preventing the policy from becoming too deterministic too quickly. In reinforcement learning, policies typically use probability distributions to choose actions. The entropy term added to the loss function measures the "randomness" of these distributions. Higher entropy means a more uncertain policy and more actions explored; lower entropy indicates higher confidence in a particular choice. By including an entropy term with an adjustable coefficient in the loss function, the optimization process can strike a balance between utilizing known good actions and exploring new ones.
[0075] Entropy can also mitigate the premature convergence of a policy to a suboptimal state. Without entropy, a policy might quickly reduce the probability of actions to almost zero, allocating them to actions that initially seem bad but may yield better rewards in the long run. For example, when an agent needs to jump over an obstacle, a deterministic policy might repeatedly fail by jumping too early. With entropy, the policy retains a portion of the probability of jumping later, allowing it to discover the correct timing.
[0076] Finally, iteratively execute steps S303-S304 until the training termination condition is met. The training termination condition is when the average round reward curve changes with the number of training steps and the reward no longer increases significantly, indicating that the policy has learned the highest level it can learn. At this point, the policy network obtained from the last training is used as the intelligent control network to perform online control.
[0077] For example, the tracking error is less than 0.1 arcminutes, and the root mean square error of stiffness realization is less than 2N.
[0078] In some embodiments, step S103 may include, but is not limited to, steps S401 to S403.
[0079] Step S401: In each control cycle, the current state vectors of the dual motors and the output shaft are acquired based on the state space.
[0080] Step S402: Input the current state vector into the policy network for forward propagation to obtain the optimal action.
[0081] Step S403: Update the optimal motor torque command of the previous control cycle according to the optimal action to obtain the updated optimal motor torque command.
[0082] In this embodiment, in each control cycle, the first angle of the first motor is acquired based on the motor encoder reading. i 1 and first speed The second angle of the second motor i 2 and the second speed The third angle of the output shaft i l With the third speed External interaction forces are collected through a torque sensor. F ext ; Receive the desired stiffness output from the task planner K d Expected perspective i ref and expected speed To obtain the current state vector s t .
[0083] For example, the angles and velocities of the two motors during operation are read by motor encoders installed on the first and second motors, respectively; the angles and velocities on the output shaft are read by an encoder installed on the output shaft; and the external interaction forces acting on the joint are measured by a torque sensor installed on the joint. Simultaneously, the desired task command is received, and the current state vector is constructed according to the parameter order in the predefined state space. s t .
[0084] The current state vector s t Input the deployed intelligent control network π φ The policy network performs forward propagation and outputs the optimal action. a t =[ △T 1, △T [2], wherein the optimal action includes a first optimal torque increment action and a second optimal torque increment action. The first optimal torque increment action is used to drive the first motor with the torque increment required in the current control cycle, and the second optimal torque increment action and the second optimal torque increment action are used to drive the second motor with the torque increment required in the current control cycle.
[0085] The optimal motor torque command for each of the two motors is updated based on the optimal action corresponding to each motor. The incremental update is a smooth evolution based on the optimal motor torque command of the previous cycle. After receiving the optimal action, the optimal motor torque command for the current control cycle is obtained by algebraically adding the optimal motor torque command of the previous control cycle to the optimal action.
[0086] In some embodiments, step S403 may include, but is not limited to, steps S501 to S502.
[0087] Step S501: Update the first motor torque command of the previous control cycle according to the first optimal torque increment action to obtain the updated first motor torque command.
[0088] Step S502: Update the second motor torque command of the previous control cycle according to the second optimal torque increment action to obtain the updated second motor torque command.
[0089] In this embodiment, the update formulas for the first motor torque command and the second motor torque command are as follows: (8); in, T 1,cmd For the updated first motor torque command, T 2,cmd For the updated second motor torque command, T 1,prev This is the first motor torque command of the previous control cycle. T 2,prev This is the second motor torque command from the previous control cycle.
[0090] Will T 1,cmd and T 2,cmd The current servo driver is sent to the two motors to drive them to run. Then the system enters the next control cycle and repeats steps S501 and S502.
[0091] The following is a detailed description and explanation of the solutions in the embodiments of the present invention, using specific application examples: Reference Figure 4 In this application embodiment, the knee joint of a humanoid robot is used as an example. During dynamic movements such as walking, the knee joint of a humanoid robot needs to adjust its joint stiffness in real time according to the gait stage and the external environment (such as ground reaction force).
[0092] First, the actual angle and angular velocity of the knee joint, the angle and angular velocity of motor 1 and motor 2, the estimated external interaction force, the desired stiffness command at the current moment, and the desired angle and desired velocity of the knee joint of the output shaft are collected.
[0093] For example, the joint encoder measures the actual angle of the load shaft. i l =12°, actual angular velocity =3° / s, measured by encoder 1 of motor. i 1 = 14° =3.5° / s, measured by encoder 2 of motor. i 2 = 10°, second velocity =2.5° / s, measured by the joint torque sensor F ext =350Nm. Desired angle i ref =10°, desired angular velocity =0° / s, desired stiffness K d =400 Nm / rad. State vectors are constructed after unifying to radians. s t =[0.2094, 0.0524, 0.2443, 0.0611, 0.1745, 0.0436, 350,400, 0.1745, 0].
[0094] Assume the intelligent variable stiffness reward function for the knee joint during the humanoid robot's walking gait on flat ground is set as follows: (9); in, r success1 For a successful item, r success2 For the item that has fallen, oh 5 represents the weighting coefficient for the success item. oh 6 represents the weighting coefficient for the fall item.
[0095] It is understandable that the variable stiffness reward function can set the weight ratios of various items according to actual needs. In this embodiment, the stiffness requirement is relatively large during operation, so the weight values are increased.
[0096] For example, the specific calculation and weights of the variable stiffness reward function are shown in Table 1.
[0097] Table 1 Reward Function Settings
[0098] in, △T 1, △T 2. The torque range is -0.6 Nm to 0.6 Nm, with a maximum torque of 20 Nm. F ext The data comes from sensor data and ranges from 0 to 1000 Nm.
[0099] The state vector s t Input the pre-trained policy network π φ ( a t | s tNetwork output actions a t =[ △T 1, △T 2]=[0.25, [0.12]. Check if the amplitude of the movement is within the preset range (-0.6Nm to 0.6Nm). If it is within the range, no action is required. If it exceeds the range, truncate to the boundary value.
[0100] Update the motor torque command based on the output action, and read the preset value from memory. T 1,prev =8.6 Nm, T 2,prev =8.1Nm, the output torque is calculated according to formula (8): (9).
[0101] Finally T 1,cmd =8.85Nm and T 2,cmd =7.98Nm is sent to the motor servo driver to drive the knee joint to perform movement.
[0102] In summary, the dual-motor seamless variable stiffness control method based on deep reinforcement learning provided in this application enables the agent to learn to automatically adjust the motor torque under any given stiffness command through the stiffness realization term in the reward function, so that the system exhibits the corresponding stiffness characteristics. The strategy obtained through reinforcement learning training is the optimal solution for long-term cumulative rewards, which can achieve the best balance between tracking accuracy, stiffness compliance and energy efficiency. It can be applied to the field of intelligent robots such as humanoid robots, humanoid robots, bipedal robots and multi-legged robots. It overcomes the problems of current variable stiffness control methods that rely on precise mathematical models, which leads to the need to design complex online optimization or parameter tuning, and it is difficult to achieve optimal performance in unknown or dynamically changing environments. It can autonomously learn control performance without the need for manual design of complex stiffness mapping or assignment laws.
[0103] Reference Figure 5 This application also provides a dual-motor seamless variable stiffness control system based on deep reinforcement learning, which can implement the above method. The system includes: The Markov decision module is used to model the dual-motor variable stiffness control task as a Markov decision process, obtaining the state space, action space and variable stiffness reward function. The state space is a set of the desired state of the task, the interaction state with the environment, and the motion states of the dual motors and the output shaft.
[0104] The policy network training module is used to perform proximal policy optimization on the policy network through state space, action space, and variable stiffness reward function, so that the policy network learns to generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thus obtaining an intelligent control network.
[0105] The control decision module is used to collect the current state vector of the two motors in each control cycle, input the current state vector into the intelligent control network for control decision-making, and obtain the optimal motor torque command.
[0106] The motor drive module is used to drive two motors to perform variable stiffness coordinated motion according to the optimal motor torque command.
[0107] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0108] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0109] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0110] Reference Figure 6 , Figure 6 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0111] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the methods described in the embodiments of this application.
[0112] The input / output interface 903 is used to implement information input and output.
[0113] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0114] Bus 905 transmits information between various components of the device, such as processor 901, memory 902, input / output interface 903, and communication interface 904.
[0115] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0116] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0117] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0118] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0119] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0120] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0121] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0122] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0123] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0124] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0125] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0126] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0127] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0128] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0129] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0130] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0131] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A seamless variable stiffness control method for dual motors based on deep reinforcement learning, characterized in that, The dual-motor system includes a first motor and a second motor, which are coupled together to output to a single output shaft. The method includes the following steps: The dual-motor variable stiffness control task is modeled as a Markov decision process, resulting in a state space, an action space, and a variable stiffness reward function. The state space is a set of the desired state of the task, the interaction state with the environment, and the motion states of the dual motors and the output shaft. By performing proximal policy optimization on the policy network using the state space, the action space, and the variable stiffness reward function, the policy network learns to generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thus obtaining an intelligent control network. In each control cycle, the current state vector of the dual motors is collected, and the current state vector is input into the intelligent control network for control decision-making to obtain the optimal motor torque command; The dual motors are driven to perform variable stiffness coordinated motion according to the optimal motor torque command.
2. The method according to claim 1, characterized in that, The state space is a set of ten-tuple data. The expected state of the task includes expected stiffness, expected angle and expected speed. The environmental interaction state is an external interaction force. The motion state includes the first angle and first speed of the first motor, the second angle and second speed of the second motor, and the third angle and third speed of the output shaft. The motion space includes the first torque increment of the first motor and the second torque increment of the second motor; The variable stiffness reward function is a multi-objective composite reward function, including a trajectory tracking term, a stiffness realization term, an energy benefit term, and a smoothing performance term. The trajectory tracking term is used to penalize trajectory tracking errors, the stiffness realization term is used to control the relationship between force and displacement, the energy benefit term is used to minimize energy consumption, and the smoothing performance term is used to penalize incremental abrupt changes.
3. The method according to claim 2, characterized in that, The process of obtaining the variable stiffness reward function includes the following steps: The trajectory error is quantized based on the first difference between the third angle and the desired angle and the second difference between the third velocity and the desired velocity to obtain the trajectory tracking term; A virtual spring simulation is performed based on the first difference and the desired stiffness to obtain a stiffness model. The spring simulation error is then quantified based on the external interaction force and the stiffness model to obtain the stiffness realization term. The energy efficiency term is obtained by performing comprehensive amplitude quantization based on the first torque increment and the second torque increment. Torque increment quantization is performed based on the first torque increment and the second torque increment to obtain the smoothing performance term; The variable stiffness reward function is obtained by weighting the trajectory tracking term, the stiffness realization term, the energy benefit term, and the smoothness performance term.
4. The method according to claim 3, characterized in that, The expression for the variable stiffness reward function is: ; in, α , β , γ , δ , η These are the weighting coefficients for each item. r t As a reward value, θ 1 represents the first angle. For the first speed, θ 2 represents the second angle. The second speed, θ l For the third angle, The third velocity, F ext For the external interaction force, K d For the desired stiffness, θ ref For the desired angle, For the desired speed, △T 1 represents the first torque increment. △T 2 represents the second torque increment.
5. The method according to claim 1, characterized in that, The process of optimizing the policy network through the state space, the action space, and the variable stiffness reward function, enabling the policy network to learn and generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thus obtaining an intelligent control network, includes the following steps: Initialize the policy network and interact with the environment through the initialized policy network to obtain training data, wherein the training data includes state data, action data and reward data; Based on the training data, a generalized dominance estimate is performed on the policy network to obtain the dominance value; The training data is randomly sampled based on a preset first quantity to obtain sample data; The policy loss function is determined based on the advantage value, and the policy network is optimized by limiting the update magnitude using the policy loss function based on the sample data to obtain the optimized policy network. Then, the process returns to the step of randomly sampling the training data based on a preset first number to obtain sample data, until the training termination condition is met. Then, the last obtained policy network is confirmed as the intelligent control network.
6. The method according to claim 2, characterized in that, The process of acquiring the current state vector of the dual motors in each control cycle, inputting the current state vector into the strategy network for control decision-making, and obtaining the optimal motor torque command includes the following steps: In each control cycle, the current state vectors of the dual motors and the output shaft are acquired according to the state space; The current state vector is input into the policy network for forward propagation to obtain the optimal action, wherein the optimal action includes a first optimal torque increment action and a second optimal torque increment action. The optimal motor torque command from the previous control cycle is updated based on the optimal action to obtain the updated optimal motor torque command.
7. The method according to claim 6, characterized in that, The optimal motor torque command includes a first motor torque command and a second motor torque command. Updating the optimal motor torque command of the previous control cycle based on the optimal action to obtain the updated optimal motor torque command includes the following steps: The first motor torque command of the previous control cycle is updated according to the first optimal torque increment action to obtain the updated first motor torque command. The second motor torque command of the previous control cycle is updated according to the second optimal torque increment action to obtain the updated second motor torque command.
8. A seamless variable stiffness control system for dual motors based on deep reinforcement learning, characterized in that, The system is used to implement the method according to any one of claims 1 to 7, the system comprising: The Markov decision module is used to model the dual-motor variable stiffness control task as a Markov decision process to obtain the state space, action space and variable stiffness reward function. The state space is a set of the task expectation state, the environmental interaction state and the motion state of the dual motors and the output shaft. The strategy network training module is used to perform proximal policy optimization on the strategy network through the state space, the action space and the variable stiffness reward function, so that the strategy network learns to generate motor torque commands that achieve variable stiffness behavior balance according to the environment, thereby obtaining an intelligent control network. The control decision module is used to collect the current state vector of the dual motors in each control cycle, input the current state vector into the intelligent control network for control decision-making, and obtain the optimal motor torque command. The motor drive module is used to drive the two motors to perform variable stiffness coordinated motion according to the optimal motor torque command.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.