A load adaptive control method for a dual-wheel dual-arm carrying robot

CN122807924APending Publication Date: 2026-09-25HARBIN INST OF TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611237324.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-14
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

实际搬运物体的质量、质心位置和夹持偏移通常难以预先准确获得,双臂运动还会持续改变整机质心和底盘平衡点,因而增加模型建立、负载参数辨识和实时控制的难度

Benefits of technology

本申请将双臂关节控制与双轮底盘轮端控制解耦,使策略网络的机器人驱动动作仅包括左驱动轮和右驱动轮的轮端控制动作,从而减少策略动作空间的维度,并有助于降低底盘平衡任务与双臂搬运任务之间的奖励耦合。双臂关节状态仍作为策略网络的观测信息,使策略网络能够根据双臂姿态变化调整轮端控制输出。策略网络利用机体姿态、轮端运动、历史控制动作和双臂关节状态构成的历史状态时序信息,可以从机器人连续动态响应中感知与未知负载影响相关的动态特征。训练阶段由价值网络接收末端负载质量、负载质心偏移和整机质心前向偏移等特权信息,以辅助策略网络更新。在线控制阶段仅使用包括必要历史状态输入部分在内的策略网络,全部特权信息均不作为策略网络的在线输入,从而降低对额外负载测量装置和在线参数辨识的依赖。策略网络输出的左右轮端控制动作经映射后由轮端执行器跟踪执行,可以形成状态采集、历史观测构建、策略推理、轮端执行和状态更新的闭环控制过程。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807924A_ABST
    Figure CN122807924A_ABST
Patent Text Reader

Abstract

The application discloses a load adaptive control method of a double-wheel double-arm carrying robot, and belongs to the technical field of reinforcement learning control of wheeled mobile robots. The double-arm joint control and the double-wheel chassis wheel end control are decoupled, so that the strategy network only outputs wheel end control actions of the left driving wheel and the right driving wheel; the historical state time sequence information composed of the body posture, the wheel end motion, the historical control action and the double-arm joint state is input into the strategy network, and the privileged information such as the end load mass, the load mass center offset and the whole machine mass center forward offset is input into the value network in the training stage; after the training under different double-arm postures and end load conditions is completed, only the strategy network constructed based on the historical state information is used for online control, and all the privileged information is not used as the online input of the strategy network. The method helps to reduce the action dimension and the reward coupling of the wheel-arm joint training, and performs the chassis balancing and speed tracking control without directly measuring the load mass and the mass center position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wheeled mobile robot control technology, and in particular to a load adaptive control method for a two-wheeled, two-arm handling robot. Background Technology

[0002] Two-wheeled self-balancing robots are characterized by their compact structure, flexible movement, and small turning radius. By mounting a dual-arm handling mechanism on a two-wheeled self-balancing mobile chassis, the robot can simultaneously perform movement and object clamping and handling, making it suitable for indoor object transport scenarios. However, due to the underactuated nature of the two-wheeled robot's support relationship with the ground, its posture and balance are easily affected by the overall center of gravity position, equivalent inertia, external disturbances, and changes in motion commands. Changes in the posture of the dual arms, as well as changes in the end-effector load mass and center of gravity position, will further alter the overall balance point.

[0003] Existing technologies employ deep reinforcement learning methods for variable load balancing control of biwheeled legged robots. For example, Chinese invention patent application CN119396197A (publication date February 7, 2025) discloses a "Variable Load Balancing Control Method, Balancing Control System, and Experimental Platform for Biwheeled Legged Robots." This approach uses domain randomization to change parameters such as robot mass, center of mass position, ground friction, and external disturbances in a physical simulation environment, and employs an actor-critic reinforcement learning framework to train the control strategy. Its motion space includes both leg joint control actions and left and right drive wheel control actions. The trained policy network is then transformed and deployed into the actual robot to achieve posture balance and motion control under variable load conditions.

[0004] The aforementioned approach enhances the adaptability of bipedal robots to changes in mass and environmental disturbances. However, when this control method is directly extended to handling robots with dual arms and end effectors, the policy network needs to simultaneously handle chassis balancing, velocity tracking, arm posture, and end effector handling tasks. Wheel-end movements, along with multiple arm joint movements, constitute the policy motion space, increasing its dimensionality. Furthermore, reward coupling may occur between the chassis balancing objective and the end effector task objective, further increasing the complexity of training and parameter tuning.

[0005] Other existing technologies establish kinematic and dynamic models for the collaborative handling of loads by dual-arm robots. For example, Chinese invention patent application CN117301064A (publication date December 29, 2023, authorization announcement number CN117301064B) discloses "A Safety Collaborative Control Method for Dual-Arm Robots Based on Fixed-Time Convergence". This technology uses a fixed-time convergence sliding mode control method to design a position controller, and also designs an internal force controller based on the clamping internal force at the end of the dual arms. The position control torque and the internal force control torque are superimposed to achieve synchronous control of the load pose and clamping force.

[0006] This type of dual-arm cooperative control method can control the end-effector trajectory and clamping internal forces of the two arms, but it usually requires establishing a dynamic model of the robot and the load, and using information such as load state, end-effector force, or clamping internal forces to complete the control calculations. When this method is used for a two-wheeled, dual-arm handling robot, the dynamic balance of the two-wheeled chassis also needs to be addressed. The mass, center of gravity position, and clamping offset of the actual object being handled are usually difficult to obtain accurately in advance, and the movement of the two arms will continuously change the center of gravity of the entire machine and the balance point of the chassis, thus increasing the difficulty of model building, load parameter identification, and real-time control.

[0007] Therefore, how to reduce the motion space dimension and reward design complexity of reinforcement learning strategies under the condition that the control of the two-arm joints and the control of the two-wheel chassis work together, and how to achieve the attitude balance and speed tracking control of the two-wheel chassis under unknown load conditions without directly measuring the end load mass and center of gravity position, has become an urgent technical problem to be solved. Summary of the Invention

[0008] The main objective of this invention is to provide a load adaptive control method for a dual-wheel, dual-arm handling robot. This method aims to reduce the complexity of reinforcement learning training by coordinating the control of the dual-arm joints and the dual-wheel chassis, and to achieve attitude balance and speed tracking of the dual-wheel chassis without relying on direct measurement of the end-load mass and center of mass position.

[0009] To achieve the above objectives, this invention proposes a load adaptive control method for a dual-wheel, dual-arm handling robot, comprising: The wheel end control of the dual-wheel chassis of the dual-wheel dual-arm handling robot is decoupled from the joint control of the dual arms. The joints of the dual arms are controlled by a joint control module independent of the strategy network. The control actions output by the strategy network only include the wheel end control actions corresponding to the left drive wheel and the right drive wheel respectively. An asymmetric actor-critic reinforcement learning framework is constructed, comprising a policy network and a value network. The input to the policy network includes the current speed command and historical state temporal information, which includes the robot's posture state, wheel-end motion state, historical control actions, and the states of both arm joints. The input to the value network includes the input to the policy network and privileged information obtained from the simulation environment, which includes end-effector load parameter information and robot global center of mass related state information. Under different dual-arm postures and end-effector load conditions, the policy network and the value network are trained using a reinforcement learning algorithm. The privileged information is used to evaluate the state value of the value network and assist in updating the policy network. After training, the policy network is used only for the online control of the dual-wheel, dual-arm handling robot. The current speed command and the historical state timing information are collected online and input into the policy network, but the privileged information is not input into the policy network. The policy network outputs the wheel-end control action, maps the wheel-end control action to the target angular velocities of the left and right drive wheels, and the wheel-end actuator tracks and executes it.

[0010] Preferably, the reinforcement learning algorithm is a proximal policy optimization algorithm.

[0011] Preferably, during the training process, multiple parallel simulation environments are constructed, and training samples of the policy network and the value network are collected simultaneously using the multiple parallel simulation environments.

[0012] Preferably, during training, the target joint posture of both arms, the end-load mass, and the position of the end-load centroid are randomly set.

[0013] Preferably, the historical state time sequence information is formed by stacking state information from 10 consecutive control moments, and the input to the policy network is normalized before being input into the policy network.

[0014] Preferably, the body attitude state includes the fuselage pitch angle and the angular velocity of the fuselage around the y-axis and z-axis in the chassis base coordinate system; the wheel end motion state includes the rotational speed of the left drive wheel and the right drive wheel; the historical control action is the wheel end control action output by the strategy network at the previous control moment; and the dual arm joint state includes the position and velocity of the shoulder joints on both sides.

[0015] Preferably, the privileged information includes the end-load mass, the spatial offset of the end-load center of mass relative to the preset clamping center of the end-grip mechanism in the chassis base coordinate system, the forward offset of the robot's overall center of mass relative to the chassis base coordinate system, and the robot base motion state information, wherein the robot base motion state information includes the base forward velocity.

[0016] Preferably, the wheel-end control action is mapped to the target angular velocity of the left drive wheel and the right drive wheel through a gain coefficient, the target angular velocity is tracked by the wheel-end speed servo actuator, and the wheel speed range and output torque range of the left drive wheel and the right drive wheel are limited in the simulation environment according to the rated speed and rated output torque of the wheel-end actuator.

[0017] Preferably, the reward function used to train the policy network and the value network includes a task tracking term, a termination penalty term, and a motion constraint term.

[0018] Preferably, the reward function is:

[0019] Where r is the total reward at the current control moment. For linear velocity tracking rewards, For yaw rate tracking rewards, For the termination penalty item, These are the constraints on the changes in wheel-end control actions at adjacent control moments. For wheel end motion constraints under zero speed command, For wheel speed and acceleration constraints, These are the angular velocity constraints for the aircraft in the roll and pitch directions. , , , , , and This represents the weight of the corresponding reward item.

[0020] Preferably, the fall of the dual-wheeled, dual-arm handling robot or the contact of either of the two arms with the ground is determined as an unsafe state, and the current training round ends based on the unsafe state, and the termination penalty is applied according to the unsafe state.

[0021] Preferably, during the training round, the forward velocity command and the yaw rate command are resampled according to a preset resampling period, and the forward velocity command and the yaw rate command are set to zero in the training samples selected according to a preset ratio; the length of the training round is preset to cover at least two velocity command changes and two arm attitude changes.

[0022] Preferably, the online control is executed cyclically in the order of state acquisition, construction of historical state time sequence information, policy network reasoning, round-end execution, and state update.

[0023] Preferably, the input encoding part of the policy network uses any one of the following to encode the historical state temporal information: a basic recurrent neural network, a long short-term memory network, a gated recurrent unit, a temporal convolutional network, or a transformer network. The above technical solution has the following advantages: This application decouples the bi-arm joint control from the wheel-end control of the bi-wheel chassis, enabling the robot's driving actions in the policy network to include only the wheel-end control actions of the left and right drive wheels. This reduces the dimensionality of the policy action space and helps to reduce the reward coupling between the chassis balancing task and the bi-arm handling task. The bi-arm joint states are still used as observation information for the policy network, allowing it to adjust the wheel-end control outputs based on changes in bi-arm posture. The policy network utilizes historical state temporal information composed of body posture, wheel-end motion, historical control actions, and bi-arm joint states to perceive dynamic characteristics related to the influence of unknown loads from the robot's continuous dynamic response. During the training phase, the value network receives privileged information such as end-load mass, load centroid offset, and overall machine centroid forward offset to assist in policy network updates. In the online control phase, only the policy network, including the necessary historical state inputs, is used; all privileged information is not used as online input to the policy network, thus reducing reliance on additional load measurement devices and online parameter identification. The left and right wheel-end control actions output by the policy network are mapped and then tracked and executed by the wheel-end actuators, forming a closed-loop control process of state acquisition, historical observation construction, policy reasoning, wheel-end execution, and state update. Attached Figure Description

[0024] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a schematic diagram of the structure of the dual-wheel, dual-arm handling robot provided in the embodiments of this application.

[0025] Figure 2 This is a schematic diagram comparing the end-to-end wheel-arm joint control method with the wheel-arm decoupled control method of this application.

[0026] Figure 3 A schematic diagram of an asymmetric actor-critic training framework for adaptive unknown load provided in an embodiment of this application. Detailed Implementation

[0027] The technical solution of this application will be described below with reference to the accompanying drawings and embodiments. The embodiments are used to illustrate the technical principles and implementation process of this application and are not intended to limit the scope of protection of this application. Where there is no conflict, the technical features in this embodiment can be combined with each other.

[0028] Example 1 This embodiment provides a load adaptive control method for a two-wheeled, dual-arm handling robot. This method is suitable for moving and transporting small or medium-sized objects in indoor environments, and is used to perform attitude balancing and speed tracking control on a two-wheeled self-balancing mobile chassis under conditions of varying arm postures and unknown end-effector load parameters.

[0029] like Figure 1As shown, the dual-wheel, dual-arm handling robot includes a left drive wheel 1, a right drive wheel 2, a dual-wheel self-balancing mobile chassis 3, a body torso 4, a left shoulder joint 5, a right shoulder joint 6, left and right robotic arms 7, and an end effector gripping mechanism 8.

[0030] The left drive wheel 1 and the right drive wheel 2 are respectively located on both sides of the dual-wheel self-balancing mobile chassis 3. The body 4 is located on the upper part of the dual-wheel self-balancing mobile chassis 3. The left and right robotic arms 7 are connected to the body 4 via the left shoulder joint 5 and the right shoulder joint 6, respectively. The end effector 8 is located at the end of the left and right robotic arms 7 and is used to grip the object to be transported from both sides.

[0031] In this embodiment, the origin of the chassis base coordinate system is set at the midpoint of the line connecting the axles of the left drive wheel 1 and the right drive wheel 2. The x-axis is along the forward direction of the robot, the y-axis is along the lateral direction of the robot, and the z-axis is vertically upward. The preset clamping center of the corresponding end gripping mechanism serves as the reference point for the spatial offset of the end load centroid.

[0032] The robot's body 4 houses a controller, power module, inertial measurement unit, and communication lines for executing control programs. The left drive wheel 1 and right drive wheel 2 are each driven by their respective wheel-end motors. The left shoulder joint 5 and right shoulder joint 6 are each driven by their respective joint actuators. The sensors, wheel-end actuators, and joint actuators interact with the controller. The controller acquires the robot's motion state and sends control commands to the wheel-end actuators and joint actuators.

[0033] The wheel-end actuator includes a wheel-end motor and its speed servo drive unit; the wheel-end motor generates driving torque, and the speed servo drive unit controls the operation of the wheel-end motor according to the target angular velocity. The wheel-end speed servo actuator referred to below is a wheel-end actuator with the aforementioned speed closed-loop control function.

[0034] When the robot performs a transport task, it first moves to the vicinity of the object to be transported. The end effector 8 grips the object from both sides. The left and right robotic arms 7 are raised under the drive of the left shoulder joint 5 and the right shoulder joint 6, causing the gripped object to leave the support surface. Subsequently, the left drive wheel 1 and the right drive wheel 2 drive the robot to move towards the target position. After the robot reaches the target position, the left and right robotic arms 7 lower the gripped object, and the end effector 8 releases the object.

[0035] During the aforementioned transport process, the changes in the posture of the left and right robotic arms 7, as well as the changes in the end-effector load mass and center of mass position, will alter the robot's overall center of mass position, equivalent inertia, and equilibrium point. Since the mass and center of mass position of the object to be transported are usually not accurately obtained in advance during actual operation, this embodiment uses the end-effector load mass, load center of mass spatial offset, and overall center of mass forward offset as privileged information during the training phase. All privileged information is not used as input to the online policy network; instead, it reflects the impact of unknown loads on the chassis's dynamic response through historical changes in the robot's measurable states.

[0036] This embodiment employs a control method that decouples the dual-arm joint control from the dual-wheel chassis wheel-end control. The dual-arm joint control is performed by a dual-arm joint control module independent of the strategy network. This dual-arm joint control module can be implemented by a high-level task planning module, a low-level joint controller, or a combination of both. It is used to control the left and right robotic arms 7 to perform lifting, holding, and lowering actions, and to cooperate with the end effector clamping mechanism 8 to complete clamping and releasing actions.

[0037] The reinforcement learning policy network does not output control actions for the left shoulder joint 5, right shoulder joint 6, or end effector 8. The control actions output by the policy network only include the wheel-end control actions corresponding to the left drive wheel 1 and right drive wheel 2, respectively. Thus, the policy network mainly learns the attitude balance and speed tracking laws of the dual-wheel self-balancing mobile chassis 3 under different dual-arm postures and different end effector load conditions.

[0038] like Figure 2 As shown, in the end-to-end joint control method for the wheel arms, the reinforcement learning policy needs to simultaneously output wheel-end movements and joint movements, and simultaneously handle chassis balance, speed tracking, joint motion, and grasping tasks. Reward-objective coupling can easily occur between different tasks. This embodiment removes the joint movements of both arms from the reinforcement learning action space, but still uses the joint states of both arms as the observation input to the policy network. This reduces the action space of the policy network while allowing it to understand the impact of changes in arm posture on chassis balance.

[0039] The load adaptive control method in this embodiment includes a reinforcement learning training process and an online control process.

[0040] In the reinforcement learning training process, a physical simulation environment for a two-wheeled, two-arm handling robot is first established. The physical simulation environment includes a two-wheeled self-balancing mobile chassis 3, a body 4, left and right robotic arms 7, an end effector 8, and corresponding joints and wheel-end actuators, all corresponding to the actual robot. The simulation environment receives wheel-end control actions and two-arm joint control commands, and outputs the robot's state, reward feedback, and the state at the next moment.

[0041] The simulation environment limits the wheel speed range and output torque range of the left drive wheel 1 and right drive wheel 2 according to the rated speed and rated output torque of the actual wheel-end actuators. This limitation ensures that the wheel-end movements generated during training are within the range that the actual wheel-end actuators can track.

[0042] like Figure 3 As shown, this embodiment constructs an asymmetric actor-critic reinforcement learning framework comprising a policy network (Actor network) and a value network (Critic network). The policy network only receives observation information obtainable during the actual online operation of the robot. The value network, based on the observations of the policy network, further receives privileged information obtainable in the physical simulation environment.

[0043] When a temporal feature extraction network is used to encode historical state temporal information, the temporal feature extraction network constitutes the input encoding part of the policy network. After training, it is used together with the policy network for online control, but is not used as an online control network independent of the policy network.

[0044] In this embodiment, the policy network at the control time The input observations are represented as follows:

[0045] in, Indicates the aircraft's pitch angle. This indicates the fuselage along the chassis base coordinate system. direction and Angular velocity in the direction. This indicates the wheel speeds of the left drive wheel 1 and the right drive wheel 2. This represents the wheel-end control action output by the policy network at the previous control moment. Indicates the position of the joints in both arms. This indicates the speed of the joints in both arms. This indicates the current speed command. The current speed command includes the forward speed command and the yaw rate command.

[0046] This represents a stacked history of 10 consecutive control moments, including fuselage pitch angle, fuselage angular velocity, wheel speed, historical control actions, arm joint positions, and arm joint velocities. The input data for each policy network is normalized before being input into the policy network.

[0047] Where t is the discrete control time index, taken as a positive integer; the superscript π indicates the strategy network observation, the superscript b indicates the chassis base coordinate system, and the subscript p indicates the pitch direction. The unit of fuselage pitch angle is radians, the unit of fuselage angular velocity and wheel speed is radians per second, the unit of boom joint position is radians, and the unit of boom joint velocity is radians per second; the current speed command is a two-dimensional vector composed of forward linear velocity command and yaw angular velocity command, with units of meters per second and radians per second, respectively. The left and right wheel speeds, historical wheel end control actions, boom joint positions, and boom joint velocities are all arranged in a fixed channel order with the left side in front and the right side behind.

[0048] H 10 (·) indicates concatenating the measurable state vectors from 10 consecutive control moments in chronological order. Normalization uses the same set of statistics determined during the training phase, and maintains the state order and normalization rules unchanged during the online phase.

[0049] The historical states over 10 consecutive control moments can reflect the robot's dynamic changes over a period of time. Different end-effector load masses and different load center-of-gravity offsets will cause the robot to produce different pitch, angular velocity, and wheel speed changes under the same wheel-end control actions. The policy network perceives the differences in dynamic response caused by unknown loads by the correspondence between historical states and historical control actions, and adjusts the wheel-end control output for control compensation.

[0050] The joint positions and velocities of both arms are also included in the historical state. When the left and right robotic arms 7 are raised, held, or lowered, the policy network can acquire the corresponding changes in arm posture and adjust the wheel-end control actions of the left drive wheel 1 and right drive wheel 2 accordingly. The joint states of the arms are input information to the policy network, but not part of its action output.

[0051] Value network at control moment The input observations are represented as follows:

[0052] in, This indicates the quality of the end load. This indicates the spatial offset of the end load centroid relative to the preset clamping center of the corresponding end clamping mechanism in the chassis base coordinate system. This indicates the forward offset of the machine's center of gravity relative to the chassis base coordinate system. This indicates the forward velocity of the robot's base.

[0053] The end-load mass m represents the mass of the object being transported, in kilograms; the spatial offset of the end-load centroid. The spatial offset of the center of mass of the object being transported relative to the preset clamping center of the end clamping mechanism 8 is expressed in the chassis base coordinate system and is in meters; the forward offset of the center of mass of the whole machine and the forward velocity of the robot base are both scalars and are in meters and meters per second, respectively.

[0054] The end-load mass, end-load centroid spatial offset, and overall centroid forward offset are all provided by the physical simulation environment during the training phase. This information is used by the value network for state value evaluation and to assist in updating the policy network parameters. None of the privileged information mentioned above is used as input to the policy network during the online runtime phase.

[0055] During training, training samples are collected simultaneously in multiple parallel physical simulation environments. Each physical simulation environment is set on a flat ground, allowing the training process to focus on learning the effects of changes in the posture of the two arms and the end load on the two-wheel self-balancing mobile chassis 3.

[0056] In different physical simulation environments, the target posture of the two arms, the end-effector load mass, and the spatial offset of the end-effector load centroid are randomly set. The dual-arm joint control module controls the left and right robotic arms 7 according to the set target angles of the two arms. The physical simulation environment inputs the corresponding dual-arm joint states into the policy network and inputs the end-effector load mass, the spatial offset of the load centroid, and the forward offset of the overall centroid into the value network.

[0057] As feasible training conditions, the target angles of the two arms are sampled within the mechanical limits of the corresponding joints; the end-effector load mass is sampled within the range from zero to the robot's rated end-effector load; the spatial offset of the end-effector load center of mass is sampled within the allowable gripping area of ​​the end-effector gripping mechanism; and the forward velocity command and yaw rate command are sampled within the rated motion range of the dual-wheel chassis. The number of parallel simulation environments can be set according to computing resources and training stability; in this embodiment, the number of parallel simulation environments is set to 4096, and each environment independently samples the above conditions.

[0058] The velocity command resampling period, the proportion of zero-velocity training samples, the bi-arm target pose update period, and the training round length are preset before training begins and remain unchanged during the same training phase; each training round contains at least two velocity command updates and two bi-arm target pose updates.

[0059] The strategy network outputs wheel-end control actions based on the current speed command and historical states:

[0060] in, This indicates the wheel end control action corresponding to the left drive wheel 1. This indicates the wheel end control action corresponding to the right drive wheel 2.

[0061] The wheel-end control actions are mapped to target angular velocities of left drive wheel 1 and right drive wheel 2 via gain coefficients. Wheel-end velocity servo actuators track the corresponding target angular velocities and apply wheel-end driving actions to the physical simulation environment. The physical simulation environment updates the robot state based on the wheel-end driving results and returns the reward feedback and the next-moment state to the reinforcement learning framework.

[0062] This embodiment employs a proximal policy optimization algorithm to train the policy network and the value network. During training, the policy network outputs wheel-end control actions based on observations. The value network estimates the value of the current state based on the policy network's observations and privileged information. The reinforcement learning algorithm updates the policy network and the value network based on the reward feedback returned from the physical simulation environment, the next time-step state, and the value estimation results.

[0063] As one feasible training setup in this embodiment, both the policy network and the value network employ multi-layer fully connected networks. The number of hidden layer neurons in the policy network is set to 512, 256, and 128, respectively, and the number of hidden layer neurons in the value network is also set to 512, 256, and 128, respectively. The activation function for both networks is the exponential linear unit function. The proximal policy optimization algorithm uses a discount factor of 0.99, a generalized dominance estimation coefficient of 0.95, a policy pruning coefficient of 0.2, and the Adam optimizer. The learning rate is set to be greater than or equal to 1 × 10⁻⁶. -4 And less than or equal to 5 × 10 -4 The collected time trajectories are iteratively updated multiple times in small batches.

[0064] The simulation step size, control cycle, batch size, single sampling length, and training termination conditions are preset based on the robot's dynamic response and computational resources, and are kept traceable in the training record; training termination can be determined based on a preset number of training steps or the reward change of multiple consecutive evaluation rounds being lower than a preset threshold.

[0065] The reward function used for training consists of a task tracking term, a termination penalty term, and a motion constraint term, and its expression is:

[0066] Where r is the total reward at the current control moment. This represents the linear velocity tracking reward, used to evaluate the tracking performance between the robot's actual forward velocity and the forward velocity command. This represents the yaw rate tracking reward, used to evaluate the tracking performance between the robot's actual yaw rate and the yaw rate command.

[0067] This indicates the termination penalty. The current training round terminates when the robot falls over or either of the left or right robotic arms (7) touches the ground, and a termination penalty is determined based on this termination state. By setting this termination condition, the possibility of the strategy using robotic arms touching the ground to support the robot or generating other unsafe actions can be reduced.

[0068] This represents the constraint term for changes in wheel-end control actions at adjacent control moments, used to suppress abrupt changes in wheel-end control actions. This represents the wheel-end motion constraint under zero-speed command, used to suppress unnecessary forward and backward movements of the robot in its stationary equilibrium state. This represents the wheel speed acceleration constraint, used to suppress drastic changes in wheel speed. The angular velocity constraints in the roll and pitch directions are used to reduce the body's attitude oscillations.

[0069] , , , , , and These are the weights for the corresponding reward items. Each weight is preset based on the robot's structural parameters, wheel-end execution capabilities, and handling task requirements.

[0070] To improve the strategy's adaptability to different motion commands, forward velocity and yaw rate commands are resampled according to a preset resampling period during training rounds. In training samples selected according to a preset ratio, the forward velocity and yaw rate commands are set to zero to train the robot to maintain balance in place. The length of a single training round is preset to cover at least two velocity command changes and two arm posture changes, enabling the strategy network to learn chassis control patterns under different velocity commands, arm postures, and load conditions during continuous transport.

[0071] After training, only the policy network, which takes the current speed command and historical state timing information as input, is retained for online robot control. The value network and all privileged information, such as end-effector mass, load centroid spatial offset, and overall centroid forward offset, are not used in online operation.

[0072] During online control, the controller first collects the fuselage attitude status output by the inertial measurement unit, the wheel speed fed back by the wheel-end motor, the position and speed of the double-arm joints fed back by the double-arm joint actuator, as well as the current speed command and the wheel-end control action at the previous control moment.

[0073] The controller updates the historical states for 10 consecutive control moments using the same state arrangement as in the training phase, and inputs the current speed command and historical states into the policy network. Before being input into the policy network, all input data are normalized according to the same set of statistics determined in the training phase, ensuring that different types of input quantities, such as fuselage attitude, angular velocity, wheel speed, arm joint states, and historical control actions, are of similar numerical magnitudes. In the online control phase, the same state arrangement order and normalization rules as in the training phase are used. After the policy network completes inference, the outputs correspond to the wheel-end control actions of the left drive wheel 1 and right drive wheel 2, respectively.

[0074] The controller maps the wheel-end control actions to target angular velocities for left drive wheel 1 and right drive wheel 2 using a gain coefficient, and sends these target angular velocities to the corresponding wheel-end speed servo actuators. The wheel-end speed servo actuators then drive left drive wheel 1 and right drive wheel 2 to rotate. The robot's state changes accordingly, and the controller re-acquires the updated state before entering the next control cycle. This forms a closed-loop control process encompassing state acquisition, historical state construction, policy network reasoning, wheel-end execution, and state updating.

[0075] During the aforementioned online control process, the dual-arm joint control module independently controls the left and right robotic arms 7 according to the handling task. The strategy network does not change the grasping and handling trajectory set by the dual-arm joint control module, but it can learn about the upper body posture changes through the state of the dual-arm joints and adjust the wheel-end control actions according to the body posture, wheel speed, and historical control response.

[0076] Since the policy network does not need to simultaneously explore the joint movements of both arms and the wheel ends, its motion space consists of the control movements of the left and right wheel ends. The reward objectives of chassis balance and speed tracking no longer directly compete with the end-effector grasping task. Meanwhile, the value network in the training phase uses load and whole-machine center-of-mass privileged information to evaluate the value of different load states, while the policy network in the online operation phase uses continuous historical states to reflect the dynamic response differences caused by unknown loads. These control relationships enable the robot to adjust the control outputs of the left drive wheel 1 and right drive wheel 2 based on its measurable states without directly measuring the end-effector load mass and load center-of-mass position.

[0077] When validating the trained strategy, different end-load masses, load centroid spatial offsets, and arm postures are selected within the training range. The robot's fall-over situation, body pitch angle error, forward velocity tracking error, and yaw rate tracking error are statistically analyzed. When comparison is required, the strategy is compared with the end-to-end joint control strategy of the wheel and arm or the control strategy that does not use privileged information from the training phase, under the same robot model, speed command, and load conditions.

[0078] Example 2 This embodiment, based on Embodiment 1, describes the method for constructing historical state temporal information. Unlike Embodiment 1, which uses a stacking method of 10 consecutive control moments, this embodiment uses a temporal feature extraction network, which serves as the input encoding part of the policy network, to encode the measurable states of the robot over a past period.

[0079] The structure of the dual-wheel, dual-arm handling robot, the decoupling relationship between the dual-arm joint control and the dual-wheel chassis wheel-end control, the asymmetric observation relationship between the policy network and the value network, and the basic processes of the training phase and the online control phase are the same as in Example 1.

[0080] At each control moment, the controller acquires the fuselage pitch angle, fuselage angular velocity, wheel speeds of left drive wheel 1 and right drive wheel 2, wheel-end control actions from the previous control moment, and the positions and speeds of both arm joints. The controller saves the measurable states of multiple consecutive control moments in chronological order and inputs the corresponding state sequences into the temporal feature extraction network.

[0081] Temporal feature extraction networks can employ basic recurrent neural networks, long short-term memory networks, gated recurrent units, temporal convolutional networks, or transformer networks. Based on the temporal correlation between consecutive states, the temporal feature extraction network outputs temporal features characterizing the robot's dynamic response. These temporal features, along with the current velocity command, serve as input to the subsequent decision-making part of the policy network.

[0082] When using a basic recurrent neural network, long short-term memory network, or gated recurrent unit, the temporal feature extraction network receives state information in the chronological order of control moments and uses its internal state to store dynamic information from previous control moments. When using a temporal convolutional network, the temporal feature extraction network performs temporal convolution processing on the state sequence of multiple consecutive control moments. When using a transformer network, the temporal feature extraction network extracts temporal features based on the correlation between different control moments in the state sequence.

[0083] The number of historical states processed by the temporal feature extraction network can be set according to the robot's control cycle and dynamic response characteristics. In addition to 10 consecutive control moments, state information from 3, 5, 8, or 15 consecutive control moments can also be used. The number of historical states determines the time range that the policy network can utilize, without changing the control relationship between the left and right wheel control actions output by the policy network.

[0084] During the reinforcement learning training phase, the policy network outputs the wheel-end control actions corresponding to the left drive wheel 1 and right drive wheel 2 based on the temporal characteristics and the current speed command. In addition to receiving the information used by the policy network, the value network also receives the end-load mass, the spatial offset of the end-load center of mass relative to the preset clamping center of the corresponding end-grip mechanism in the chassis base coordinate system, the forward offset of the overall center of mass relative to the chassis base coordinate system, and the forward velocity of the robot base.

[0085] During training, by changing the target angles of the two arms and the end-effector load conditions, the robot is made to produce different posture responses and wheel-end motion responses under the same control actions. The temporal feature extraction network perceives the influence of load changes and arm posture changes on relevant temporal features from continuous state changes. The value network uses privileged information from the training phase to evaluate state value and assists the policy network and temporal feature extraction network in updating parameters.

[0086] After training, the policy network, including the temporal feature extraction network, is retained for online control. The value network and all privileged information from the training phase are not used in online operation.

[0087] During online control, the controller periodically collects the measurable states of the robot and inputs them into the temporal feature extraction network in chronological order. The temporal feature extraction network outputs the temporal features at the current moment. The policy network outputs the control actions for the left and right wheels based on these temporal features and the current speed command.

[0088] The controller maps the control actions of the left and right wheels to the target angular velocities of the left drive wheel 1 and the right drive wheel 2, which are then tracked and executed by the corresponding wheel-end speed servo actuators. State acquisition, timing feature updates, policy network inference, and wheel-end execution are performed cyclically, thus forming an online closed-loop control.

[0089] This embodiment utilizes a temporal feature extraction network to replace the explicit fixed-frame stacking method. The policy network still does not directly receive end-load mass, load centroid spatial offset, and overall machine centroid forward offset. Instead, it extracts the unknown load influence based on the robot's body posture, wheel speed, arm joint states, and historical control responses over a period of time.

[0090] Example 3 This embodiment describes alternative implementations of the reinforcement learning algorithm based on Embodiment 1. The structure, wheel-arm decoupling control relationship, policy network observation, value network privileged observation, and online control process of the two-wheeled, two-arm handling robot are the same as in Embodiment 1.

[0091] Reinforcement learning training algorithms are not limited to proximal policy optimization algorithms. Depending on the training requirements of continuous wheel-end control actions, soft actor-critic algorithms, deep deterministic policy gradient algorithms, dual-delay deep deterministic policy gradient algorithms, dominant actor-critic algorithms, or other reinforcement learning algorithms that simultaneously satisfy the conditions of two-dimensional continuous action output and training phase added value evaluation information input can also be used.

[0092] When employing the soft actor-critic algorithm, the policy network outputs control actions for the left and right wheels based on the current speed command and historical state timing information. The critic network evaluates the corresponding states and actions using observations from the policy network and privileged information from the training phase. During training, both the policy network and the critic network are updated based on the rewards returned from the simulation environment and the next state.

[0093] When employing the deep deterministic policy gradient algorithm or the dual-delay deep deterministic policy gradient algorithm, the policy network outputs deterministic left and right wheel-end control actions. The critic network evaluates the system based on the state information and the wheel-end control actions. End-load mass, load centroid spatial offset, and overall machine centroid forward offset are only used during the training phase to aid evaluation and parameter updates.

[0094] When employing the dominant actor-critic algorithm, the policy network determines wheel-end control actions based on measurable states, while the value network estimates state values ​​using additional privileged information. After training, only the policy network is retained for online control. During online control, the policy network outputs left and right wheel-end control actions based on historical state timing information and the current speed command; all privileged information is not input into the online policy network.

[0095] Regardless of the reinforcement learning algorithm used, the control actions output by the policy network only include the wheel-end control actions of the left drive wheel 1 and right drive wheel 2. The left and right robotic arms 7 and the end effector 8 are still controlled by independent dual-arm joint control modules. The joint positions and velocities of the two arms are used as observations by the policy network to reflect the impact of arm posture changes on chassis balance. Different arm postures and end effector load conditions are set during training. The value network or the training module corresponding to value assessment can utilize privileged information such as end effector load mass, end effector load centroid spatial offset, and overall machine centroid forward offset. After training, all privileged information is not used as input to the online policy network.

[0096] For the near-end policy optimization algorithm and the dominant actor-critic algorithm, training samples are collected according to the time trajectory, and the policy network and value evaluation module are updated based on the collected time trajectory. For the soft actor-critic algorithm, the deep deterministic policy gradient algorithm, and the double-delay deep deterministic policy gradient algorithm, training samples are stored in the experience replay cache and updated according to mini-batch sampling. Regardless of the algorithm used, the policy network outputs two-dimensional continuous wheel-end control actions. The value evaluation module can receive the privileged information during the training phase. During the online phase, only the policy network is retained for online control, and all privileged information is not input into the online policy network.

[0097] Example 4 This embodiment, based on Embodiment 1, explains the expansion methods of the policy network measurable observation information and the value network privileged information.

[0098] The basic observations of the policy network include the fuselage pitch angle, fuselage angular velocity, left and right wheel speeds, historical control actions, joint positions of both arms, joint speeds of both arms, and current speed commands. When the robot's sensors and wheel-end actuators can provide the corresponding data, the policy network observations may also include the fuselage linear velocity, wheel-end current, estimated wheel-end torque, estimated joint torque, acceleration output by the inertial measurement unit, and the operating state of the end effector 8.

[0099] The linear velocity of the robot body is used to characterize the translational motion of the robot body. Wheel-end current and estimated wheel-end torque reflect the load and execution response of the left drive wheel 1 and right drive wheel 2. Joint estimated torque reflects the force changes of the left and right robotic arms 7 under different postures and load conditions. The acceleration output by the inertial measurement unit reflects changes in the robot body's motion state. The operating state of the end effector 8 distinguishes between gripping, holding, and releasing processes.

[0100] The aforementioned extended observations all pertain to the state information obtainable during the actual operation of the robot. The controller combines these observations with the historical states from Example 1 to form the input to the policy network. Each observation data is normalized before being input into the policy network according to the processing method used during the training phase.

[0101] In addition to the end-load mass, end-load centroid spatial offset, and overall centroid forward offset, the privileged information of the value network may also include the overall centroid's three-dimensional position, centroid velocity, load inertia, contact force, and ground friction coefficient. This information is provided by the physical simulation environment during the training phase to improve the value network's ability to distinguish between different load and contact states.

[0102] The extended privileged information is only used for state value evaluation or policy updates during the training phase; all privileged information is not used as input during online control of the policy network. During actual robot operation, the policy network still outputs wheel-end control actions for the left drive wheel 1 and right drive wheel 2 based on measurable historical states and current speed commands.

[0103] Example 5 This embodiment, based on Embodiment 1, explains the organization of arm movements and the process of carrying tasks during training.

[0104] The dual-arm joint control module generates the target poses of the left and right robotic arms 7 based on the handling task. The reinforcement learning policy network does not participate in the generation of the target poses of the arms, nor does it output the joint movements of the left shoulder joint 5 and the right shoulder joint 6.

[0105] At the start of training, the dual-arm joint control module controls the left and right robotic arms 7 in a preset initial posture. The physical simulation environment sets the end-effector load conditions and provides the policy network with the robot's current measurable state. The policy network outputs control actions for the left and right wheels, enabling the dual-wheel self-balancing mobile chassis 3 to maintain balance or move according to the current speed command.

[0106] During the training rounds, the dual-arm joint control module changes the target angles of the left and right robotic arms 7, causing them to perform lifting, holding, or lowering actions. The positions and velocities of the dual-arm joints change during the actions and are included as part of the policy network's historical state.

[0107] During arm attitude changes, the policy network cannot assist in chassis balance by altering arm joint movements. The policy network needs to adjust wheel-end outputs based on fuselage pitch angle, fuselage angular velocity, left and right wheel speeds, arm joint states, and historical control actions. Therefore, arm attitude changes, as observable but not directly controllable upper body state changes by the policy network, enter the chassis control process.

[0108] The physical simulation environment can also change the spatial offset of the end-load mass and load centroid relative to the preset clamping center of the corresponding end-grip mechanism in the chassis base coordinate system during different training rounds. The value network receives the corresponding load privilege information. The policy network does not receive the spatial offset of the load mass and load centroid, but instead learns the wheel-end control laws under different load conditions based on the robot's continuous dynamic response.

[0109] During online material handling, the dual-arm joint control module sequentially controls the end effector 8 to grip the object, the left and right robotic arms 7 to lift the object, maintain the handling posture, lower the object, and release the object. Throughout this process, the strategy network continuously outputs the wheel-end control actions of the left drive wheel 1 and the right drive wheel 2.

[0110] When the posture of the left and right robotic arms 7 changes, the policy network adjusts the wheel-end outputs based on the updated joint states of the arms. When the robot grips an unknown load, the policy network adjusts the wheel-end outputs based on changes in the robot's posture, wheel speed, and historical control responses. The dual-arm handling task and the chassis balancing task are decoupled through motion outputs and coordinated through state observation.

[0111] The above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit the scope of protection of this application. Without departing from the technical concept of this application, those skilled in the art can make adjustments to the number of historical states, temporal feature extraction methods, reinforcement learning algorithms, measurable observation information, training phase privileged information, reward items, and control action mapping methods based on the actual robot structure, sensor configuration, actuator performance, and handling task requirements, while satisfying the wheel-arm decoupling relationship, the two-dimensional continuous wheel-end action output relationship, and the condition that no privileged information is input during the online phase.

[0112] For example, different numbers of consecutive historical states can be used, or basic recurrent neural networks, long short-term memory networks, gated recurrent units, temporal convolutional networks, or transformer networks can be used to extract temporal features from historical states. Reinforcement learning algorithms can employ proximal policy optimization algorithms, soft actor-critic algorithms, deep deterministic policy gradient algorithms, double-delay deep deterministic policy gradient algorithms, dominant actor-critic algorithms, or other reinforcement learning algorithms that simultaneously satisfy the conditions of two-dimensional continuous wheel-end action output and training phase added value evaluation information input.

[0113] The measurable observations used by the policy network can be selected based on the robot's sensor configuration. The privileged information used by the value network can be selected based on the information available from the physical simulation environment. The selected additional information is only used to assist policy training during the training phase, and all privileged information is not used as input to the online policy network; while maintaining the input-output relationship and training conditions, it can be used for the unknown load adaptive control described in this application.

[0114] The dual-arm joint control module can employ a high-level task planning module, a low-level joint controller, or a combination of both. Wheel-end control actions can be mapped to target angular velocities or corresponding equivalent drive commands based on the control interface of the wheel-end actuators. The above adjustments do not alter the technical concept of decoupling the dual-arm joint control from the wheel-end reinforcement learning control, using the dual-arm states as observations for the policy network, using load-related privileged information during the training phase, not using any privileged information as input to the online policy network, and utilizing historically measurable states for chassis control during the online phase.

[0115] Equivalent substitutions, conventional modifications, or combinations made by those skilled in the art based on the disclosure of this application, as long as they do not depart from the essence of the technical solution of this application, are all reasonable modifications of the technical solution of this application.

Claims

1. A load adaptive control method for a dual-wheel, dual-arm handling robot, characterized in that, include: The wheel end control of the dual-wheel chassis of the dual-wheel dual-arm handling robot is decoupled from the joint control of the dual arms. The joints of the dual arms are controlled by a joint control module independent of the strategy network. The control actions output by the strategy network only include the wheel end control actions corresponding to the left drive wheel and the right drive wheel respectively. An asymmetric actor-critic reinforcement learning framework is constructed, comprising a policy network and a value network. The input to the policy network includes the current speed command and historical state temporal information, which includes the robot's posture state, wheel-end motion state, historical control actions, and the states of both arm joints. The input to the value network includes the input to the policy network and privileged information obtained from the simulation environment, which includes end-effector load parameter information and robot global center of mass related state information. Under different dual-arm postures and end-effector load conditions, the policy network and the value network are trained using a reinforcement learning algorithm. The privileged information is used to evaluate the state value of the value network and assist in updating the policy network. After training, the policy network is used only for the online control of the dual-wheel, dual-arm handling robot. The current speed command and the historical state timing information are collected online and input into the policy network, but the privileged information is not input into the policy network. The policy network outputs the wheel-end control action, maps the wheel-end control action to the target angular velocities of the left and right drive wheels, and the wheel-end actuator tracks and executes it.

2. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 1, characterized in that, The reinforcement learning algorithm is a near-end policy optimization algorithm.

3. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 2, characterized in that, During the training process, multiple parallel simulation environments are constructed, and training samples of the policy network and the value network are collected simultaneously using these multiple parallel simulation environments.

4. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 1, characterized in that, During training, the target joint posture, end-load mass, and end-load centroid position of both arms are randomly set.

5. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 1, characterized in that, The historical state time sequence information is formed by stacking the state information of 10 consecutive control moments, and the input of the policy network is normalized before being input into the policy network.

6. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 5, characterized in that, The aircraft attitude state includes the aircraft pitch angle and the angular velocity of the aircraft around the y-axis and z-axis in the chassis base coordinate system. The wheel end motion state includes the rotational speed of the left drive wheel and the right drive wheel. The historical control action is the wheel end control action output by the strategy network at the previous control moment. The dual arm joint state includes the position and velocity of the shoulder joints on both sides.

7. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 6, characterized in that, The privileged information includes the end-load mass, the spatial offset of the end-load center of mass relative to the preset clamping center of the end-grip mechanism in the chassis base coordinate system, the forward offset of the robot's overall center of mass relative to the chassis base coordinate system, and the robot base motion state information, including the robot base forward velocity.

8. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 1, characterized in that, The wheel-end control action is mapped to the target angular velocity of the left and right drive wheels through a gain coefficient. The wheel-end speed servo actuator tracks the target angular velocity, and the wheel speed range and output torque range of the left and right drive wheels are limited in the simulation environment according to the rated speed and rated output torque of the wheel-end actuator.

9. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 1, characterized in that, The reward function used to train the policy network and the value network includes a task tracking term, a termination penalty term, and a motion constraint term.

10. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 9, characterized in that, The reward function is: Where r is the total reward at the current control moment. For linear velocity tracking rewards, For yaw rate tracking rewards, For the termination penalty item, These are the constraints on the changes in wheel-end control actions at adjacent control moments. For wheel end motion constraints under zero speed command, For wheel speed and acceleration constraints, These are the angular velocity constraints for the aircraft in the roll and pitch directions. , , , , , and This represents the weight of the corresponding reward item.

11. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 10, characterized in that, If the dual-wheeled, dual-arm handling robot falls to the ground or either of its arms touches the ground, it is considered an unsafe state. The current training round ends based on the unsafe state, and the termination penalty is applied according to the unsafe state.

12. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 10, characterized in that, During the training round, the forward velocity command and the yaw rate command are resampled according to a preset resampling period, and the forward velocity command and the yaw rate command are set to zero in the training samples selected according to a preset ratio; the length of the training round is preset to cover at least two velocity command changes and two arm attitude changes.

13. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 1, characterized in that, Online control is executed cyclically in the order of status acquisition, construction of historical status time sequence information, policy network reasoning, round-end execution, and status update.

14. The load adaptive control method for a dual-wheel, dual-arm handling robot according to claim 1, characterized in that, The input encoding part of the policy network uses any one of the following to encode the historical state temporal information: basic recurrent neural network, long short-term memory network, gated recurrent unit, temporal convolutional network, or transformer network.

Citation Information

Patent Citations

  • Double-arm robot safety cooperative control method based on fixed time convergence

    CN117301064A

  • A safe cooperative control method for dual-arm robots based on fixed-time convergence

    CN117301064B

  • Variable load balance control method and balance control system of double-wheel-foot type robot and experimental platform

    CN119396197A