A Reinforcement Learning-Based Cooperative Control Method for Dynamic Loads on Carbon Fiber Joints of Robots
By employing a reinforcement learning-based dynamic load collaborative control method for robot carbon fiber joints, the material health status is monitored and compensated in real time, solving the performance degradation problem caused by creep and aging of carbon fiber joints in existing technologies. This achieves a balance between the long-term health reliability and short-term control performance of robot joints.
Patent Information
- Application Number
- CN202511324363.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing robot control methods, in pursuit of instantaneous task accuracy, lack effective prediction and coordinated compensation for time-varying effects of carbon fiber joints such as creep and aging, making it difficult to balance short-term control performance with long-term structural health and reliability.
A reinforcement learning-based dynamic load collaborative control method for carbon fiber joints in robots is adopted. By acquiring sensor data and using a physical observation model to calculate the health status parameters of the joint material, a reinforcement learning agent outputs collaborative control actions to generate command torque and stiffness adjustment, thereby realizing real-time monitoring and compensation of the material's health status.
It improves the control robustness under conditions of changing material properties, enables multi-dimensional adjustments based on task requirements and joint health status, extends the service life of the robot's carbon fiber joint, and ensures long-term consistency and reliability of control performance.
Smart Images

Figure CN120816501B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of humanoid robot technology, specifically to a method for dynamic load collaborative control of carbon fiber joints in robots based on reinforcement learning. Background Technology
[0002] In recent years, with the development of robotics technology towards high dynamics, high load capacity, and lightweight, carbon fiber composite materials have been increasingly widely used in robot structural design, especially in key load-bearing parts such as the limbs and joints of humanoid robots, due to their excellent properties of high specific strength and high specific modulus.
[0003] Existing technologies for controlling robots with carbon fiber composite joints often rely on idealized rigid body dynamics assumptions, simplifying robot joints into components with constant physical properties. This control method achieves good trajectory tracking performance when performing short-term, low-load tasks. However, carbon fiber reinforced polymer matrix composites inherently possess viscoelastic properties, exhibiting time-varying effects such as creep, stress relaxation, and material aging when subjected to long-term cyclic loading or continuous high-temperature working environments.
[0004] Existing control strategies typically prioritize instantaneous task execution accuracy, but lack effective online observation and feedforward compensation mechanisms for slowly accumulating end effector position deviations caused by factors such as material creep. Furthermore, control systems often fail to adequately consider the evolution of the health status of joint materials during decision-making. The generation of control laws usually does not include the goal of actively maintaining long-term material health and delaying performance degradation, which may lead to problems such as decreased control accuracy, permanent joint deformation, and reduced reliability after long-term service.
[0005] Therefore, this invention proposes a reinforcement learning-based dynamic load collaborative control method for robot carbon fiber joints to address the shortcomings of existing technologies. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a dynamic load collaborative control method for robot carbon fiber joints based on reinforcement learning. This method solves the problem that robot control methods, while pursuing instantaneous task accuracy, lack effective prediction and collaborative compensation methods for the performance degradation of carbon fiber joints caused by time-varying effects such as creep and aging, making it difficult to balance the robot's short-term control performance and long-term structural health and reliability.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a dynamic load cooperative control method for robot carbon fiber joints based on reinforcement learning, the method comprising the following steps:
[0008] S1. Acquire sensing data within the robot's carbon fiber joint, the sensing data including strain and temperature;
[0009] S2. Based on the sensing data, calculate the material physical state parameters of the joint material in the mid-term health state and the material baseline physical parameter vector of the joint material in the long-term health state through the physical state observation model.
[0010] S3. The sensing data, material physical state parameters, and material baseline physical parameter vector are used together as the state input of the reinforcement learning agent, and the reinforcement learning agent outputs a coordinated control action including torque adjustment action and physical stiffness adjustment action.
[0011] S4. Based on the coordinated control action, generate the command torque to drive the joint and the command stiffness to adjust the controllable stiffness unit within the joint, and simultaneously execute the command torque and command stiffness on the joint.
[0012] S5. Based on the execution history of the command torque and command stiffness, update the material baseline physical parameter vector in the physical state observation model, and provide the updated material baseline physical parameter vector to the physical state observation model for subsequent calculations.
[0013] Preferably, in step S1, the step of acquiring the sensing data within the robot's carbon fiber joint includes:
[0014] The strain and temperature are collected by a distributed fiber optic grating sensor array embedded inside the robot's carbon fiber joint structure.
[0015] Preferably, in step S2, the material physical state parameters for calculating the mid-term health state of the joint material using the physical state observation model include:
[0016] Creep damage measurement parameters are used to quantify the cumulative energy dissipation of materials due to creep.
[0017] Material microstructure stability parameters are used to characterize the degree of crystallinity and stability of amorphous regions in a material.
[0018] Preferably, the calculation steps for the creep damage measurement parameter include:
[0019] Based on the sensor data, the stress and creep strain rate inside the joint are calculated.
[0020] Based on the stress and creep strain rate, the creep damage measurement parameter is calculated, and the real-time change rate of the creep damage measurement parameter satisfies the following formula:
[0021] ;
[0022] In the formula, For time; The real-time rate of change of the creep damage measurement parameter; This refers to the stress inside the joint; denoted as creep strain rate.
[0023] Preferably, in step S2, the material baseline physical parameter vector of the joint material in the long-term health state includes at least the creep coefficient, time exponent and initial elastic modulus used to describe the material creep compliance.
[0024] Preferably, step S3 specifically includes:
[0025] The reinforcement learning agent is trained by maximizing the long-term cumulative value of a composite reward function, which includes:
[0026] A performance reward item used to characterize the accuracy of joint motion trajectory tracking;
[0027] Mid-term health maintenance bonus item, which is negatively correlated with the degree of degradation of the material's physical state parameters;
[0028] The long-term health gain reward term is negatively correlated with the time rate of change of the material baseline physical parameter vector.
[0029] Preferably, the controllable stiffness unit within the joint in step S4 is a magnetorheological damper, an electrorheological damper, or a piezoelectric stack actuator.
[0030] Preferably, in step S4, the step of generating the command torque to drive the joint includes:
[0031] The predicted creep displacement is superimposed on the desired trajectory to obtain the corrected desired trajectory;
[0032] The proportional-integral-differential (PI-DI) fundamental controller generates the fundamental torque based on the corrected desired trajectory.
[0033] The command torque is obtained by adding the base torque to the compensation torque generated according to the torque adjustment action.
[0034] Preferably, step S5, the step of updating the material baseline physical parameter vector within the physical state observation model, includes:
[0035] The update of the material baseline physical parameter vector is based on the following evolution model:
[0036] ;
[0037] In the formula, For time; The rate of change of the material baseline physical parameter vector; This is a preset evolution function; For the stress history of the joint; The temperature history of the joint; In time The material baseline physical parameter vector;
[0038] The evolution function is based on the Arrhenius model and is used to characterize the accelerating effect of temperature on material aging.
[0039] This invention also provides a reinforcement learning-based dynamic load cooperative control system for robot carbon fiber joints, the system comprising:
[0040] The sensor data acquisition module is used to acquire sensor data within the robot's carbon fiber joint, including strain and temperature.
[0041] The physical state observation module is used to receive the sensing data and calculate the material physical state parameters of the joint material in the mid-term health state and the material baseline physical parameter vector of the joint material in the long-term health state based on the sensing data.
[0042] The reinforcement learning decision module is used to receive the sensing data, material physical state parameters, and material baseline physical parameter vector as state inputs, and output a coordinated control action including torque adjustment action and physical stiffness adjustment action.
[0043] The collaborative control command generation and execution module is used to generate and synchronously execute the command torque for driving the joint and the command stiffness for adjusting the controllable stiffness unit within the joint, based on the collaborative control action.
[0044] The baseline parameter update module is used to update the material baseline physical parameter vector in the physical state observation module according to the execution history of the command torque and command stiffness, and provide the updated material baseline physical parameter vector to the physical state observation module.
[0045] This invention provides a reinforcement learning-based method for dynamic load cooperative control of carbon fiber joints in robots. It offers the following advantages:
[0046] 1. This invention constructs a physical state observation model to calculate in real time the material physical state parameters characterizing the mid-term health state of the joint material and the material baseline physical parameter vector characterizing the long-term health state. These parameters are used as inputs to the reinforcement learning agent. This setting enables the control decision to distinguish the source of the task execution error, such as whether it is caused by improper control instructions or by material creep, thereby performing more targeted control adjustments and improving the control robustness under the condition of material property changes.
[0047] 2. This invention generates and synchronously executes command torque and command stiffness by outputting a cooperative control action that includes torque adjustment action and physical stiffness adjustment action through a reinforcement learning agent. This cooperative control method endows the system with multi-dimensional adjustment capabilities for handling dynamic loads, enabling it to coordinate and optimize active torque compensation and passive stiffness adjustment according to different task requirements and joint health status, thereby reducing stress impact on the joint structure while ensuring task execution.
[0048] 3. Based on the execution history of command torque and command stiffness, this invention updates the material baseline physical parameter vector in the physical state observation model online. This mechanism establishes an adaptive closed-loop control model that evolves with the actual physical properties of the material, enabling the control system to track and adapt to irreversible changes in the material caused by factors such as aging and fatigue, thus ensuring the consistency and long-term effectiveness of control performance throughout the entire life cycle of the robot joint.
[0049] 4. This invention introduces a penalty term for the medium- and long-term health status of materials into the composite reward function of reinforcement learning, guiding the optimization direction of the control strategy from a single task accuracy target to a comprehensive optimal balance between task accuracy and long-term joint health. This enables the system to learn and execute material-friendly load patterns, actively avoid or mitigate operations that lead to material performance degradation, thereby effectively extending the service life of the robot's carbon fiber joints. Attached Figure Description
[0050] Figure 1 This is a structural block diagram of the robot carbon fiber joint dynamic load collaborative control system of the present invention;
[0051] Figure 2 This is a flowchart of the robot carbon fiber joint dynamic load collaborative control method of the present invention;
[0052] Figure 3 This is an internal structural block diagram of a specific embodiment of the matter observation module of the present invention;
[0053] Figure 4 This is a schematic diagram of the composite reward function of the present invention;
[0054] Figure 5 This is a schematic diagram of the command torque generation process of the present invention.
[0055] Among them, 10 is the sensor data acquisition module; 20 is the object state observation module; 30 is the reinforcement learning decision module; 40 is the cooperative control command generation and execution module; and 50 is the baseline parameter update module. Detailed Implementation
[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] See attached document Figure 1 , Figure 1 This is a structural block diagram of a robot carbon fiber joint dynamic load collaborative control system according to an embodiment of the present invention; the present invention provides a robot carbon fiber joint dynamic load collaborative control system, which may include: a sensor data acquisition module 10, a physical state observation module 20, a reinforcement learning decision module 30, a collaborative control command generation and execution module 40, and a baseline parameter update module 50;
[0058] The sensor data acquisition module 10 is configured to acquire sensor data within the robot's carbon fiber joint, including strain and temperature; the sensor data acquisition module 10 transmits the acquired sensor data to the physical state observation module 20 and the reinforcement learning decision module 30, respectively.
[0059] The physical state observation module 20 is connected to the sensing data acquisition module 10 and is used to receive sensing data. The physical state observation module 20 is configured to calculate the material physical state parameters characterizing the mid-term health state of the joint material and the material baseline physical parameter vector characterizing the long-term health state of the joint material based on the sensing data. The calculated results are transmitted to the reinforcement learning decision module 30.
[0060] The reinforcement learning decision module 30 is connected to the sensor data acquisition module 10 and the physical state observation module 20, respectively. The reinforcement learning decision module 30 is configured to take the received sensor data, material physical state parameters, and material baseline physical parameter vector as state inputs, and output a coordinated control action including torque adjustment action and physical stiffness adjustment action according to the state input. The coordinated control action is transmitted to the coordinated control instruction generation and execution module 40.
[0061] The collaborative control command generation and execution module 40 is connected to the reinforcement learning decision module 30. The collaborative control command generation and execution module 40 is configured to generate the command torque for driving the joint and the command stiffness for adjusting the controllable stiffness unit within the joint based on the received collaborative control action, and to execute the command torque and command stiffness synchronously on the joint. At the same time, the collaborative control command generation and execution module 40 transmits the execution history data of the command torque and command stiffness it generates to the baseline parameter update module 50.
[0062] The baseline parameter update module 50 is connected to the cooperative control command generation and execution module 40. The baseline parameter update module 50 is configured to update the material baseline physical parameter vector stored in the physical state observation module 20 according to the execution history of the received command torque and command stiffness. Through this update, the physical state observation module 20 will use the updated parameters in subsequent calculations, thereby forming a closed-loop regulation path.
[0063] See attached document Figure 2 , Figure 2 This is a flowchart of a dynamic load coordinated control method for a robot carbon fiber joint according to an embodiment of the present invention. The method generally includes the following steps:
[0064] Step S1: Acquire sensor data within the robot's carbon fiber joint, including strain and temperature.
[0065] Step S2: Input the sensing data into the physical state observation model, and calculate the material physical state parameters of the joint material in the mid-term health state and the material baseline physical parameter vector of the joint material in the long-term health state.
[0066] Step S3: The sensor data, material physical state parameters, and material baseline physical parameter vector are used as the state input of the reinforcement learning agent, and the reinforcement learning agent outputs a coordinated control action that includes torque adjustment action and physical stiffness adjustment action.
[0067] Step S4: Based on the coordinated control action, generate the command torque to drive the joint and adjust the command stiffness of the controllable stiffness unit within the joint, and simultaneously execute the command torque and command stiffness on the joint.
[0068] Step S5: Update the material baseline physical parameter vector in the physical state observation model based on the execution history of the command torque and command stiffness, so that the physical state observation model can use it in subsequent calculations.
[0069] The components and steps of the robot carbon fiber joint dynamic load collaborative control system and method provided by the present invention will be described in detail below.
[0070] See attached document Figure 1 Step S1 is by Figure 1The sensor data acquisition module 10 shown performs the following: In a specific embodiment, the sensor data acquisition module 10 includes a distributed fiber Bragg grating sensor array, a broadband light source, and a fiber Bragg grating demodulator; the distributed fiber Bragg grating sensor array is directly embedded into the interlayer structure of the carbon fiber reinforced polymer matrix during the manufacturing process of the robot's carbon fiber joint; the array consists of multiple optical fibers, each optical fiber having multiple fiber Bragg gratings (FBGs) with different reflection wavelengths etched on it, and these gratings are distributed along a preset path to cover key stress concentration areas and temperature-sensitive areas inside the joint;
[0071] During system operation, a broadband light source emits a broadband beam of light, which is injected into the optical fiber of the distributed fiber Bragg grating sensor array. As the beam propagates through the optical fiber, each fiber Bragg grating reflects the light of its specific center wavelength (i.e., the Bragg wavelength) back into the optical path. When the strain or temperature of the joint material changes, it will cause the period or effective refractive index of the grating to change, thereby causing the reflected Bragg wavelength to drift.
[0072] The function of the fiber Bragg grating demodulator is to detect the spectrum reflected from each fiber Bragg grating in real time and with high precision, and to accurately calculate the Bragg wavelength shift of each grating. Since the wavelength shift of a single fiber Bragg grating is affected by both strain and temperature, this embodiment uses a temperature compensation configuration in order to accurately separate these two physical quantities.
[0073] In one specific configuration, for each grating used to measure strain, a compensation grating for temperature measurement is arranged nearby; the compensation grating is encapsulated in a microcapillary unaffected by mechanical stress to ensure that it responds only to temperature changes; therefore, the fiber grating demodulator first calculates the precise temperature value at its location based on the wavelength shift of the compensation grating.
[0074] Subsequently, the fiber optic grating demodulator uses the known temperature-wavelength drift relationship to subtract the drift component caused by temperature change from the total wavelength drift of the paired strain measurement grating. After this compensation calculation, the remaining wavelength drift is the drift caused only by mechanical strain. The precise strain value at that point can be calculated using the calibrated strain-wavelength drift relationship.
[0075] Finally, the sensor data acquisition module 10 integrates the data demodulated from all grating points to generate a discretized strain field distribution map and temperature field distribution map that characterize the internal state of the entire joint. These two sets of sensor data containing spatial location information will be output as digital signals to the matter observation module 20 and the reinforcement learning decision module 30.
[0076] In other embodiments, the sensor data acquisition module 10 may also employ other types of sensors to acquire strain and temperature data. For example, a patch or sputtered piezoresistive strain gauge array can be used to measure strain, and a thermistor or thermocouple can be used in conjunction to measure temperature. Furthermore, data transmission between the sensor data acquisition module 10, the matter observation module 20, and the reinforcement learning decision module 30 can be achieved via a high-speed industrial bus (e.g., CAN bus or EtherCAT) to ensure data real-time performance and reliability.
[0077] See attached document Figure 1 and attached Figure 3 Step S2 is by Figure 1 The physical state observation module 20 shown executes the following: The physical state observation module 20 receives the sensing data provided by the sensing data acquisition module 10. Its core function is to transform the original strain and temperature data, which have relatively simple physical meanings, into a multi-dimensional parameter set that can comprehensively characterize the internal health state of the robot's carbon fiber joint material, including material physical state parameters that characterize the intermediate state and material baseline physical parameter vectors that characterize the long-term state.
[0078] In one specific embodiment, the dynamic behavior of the robot's carbon fiber joint can be described by the following equation:
[0079] ;
[0080] In the formula, , , These are the vectors of the joint's angle, angular velocity, and angular acceleration, respectively. Here is the mass inertia matrix; For Coriolis force and centrifugal force terms; This is the term related to gravity. This provides the output torque for the motor. This refers to the equivalent disturbance torque generated by the material creep effect. One object of this invention is to address... Perform precise modeling and compensation.
[0081] See attached document Figure 3 In a specific embodiment, the internal functions of the physical state observation module 20 can be divided into: a stress-strain decoupling unit, a mid-term health state calculation unit, and a long-term health state calculation unit.
[0082] The stress-strain decoupling unit receives the total strain provided by the sensing data acquisition module 10. ) and temperature ( Since the total strain is elastic strain (); ) and creep strain ( The superposition of ) Subsequent health status assessments require distinguishing between different strain components and their corresponding stresses; therefore, this unit first performs decoupling calculations. This unit reads the current material baseline physical parameter vector updated by the baseline parameter update module 50 from internal storage and obtains the real-time elastic modulus (…). Subsequently, the constitutive relations of the material are solved simultaneously by solving a system of equations (e.g., ) and creep models (e.g., the Findley creep model), which isolate the current internal stress from the total strain ( ) and creep strain ( The creep strain rate () is calculated by numerically differentiating the historical creep strain sequence. This unit will calculate the stress ( ), creep strain rate ( ) and temperature ( The data is transmitted to the mid-term health status calculation unit.
[0083] The intermediate health status calculation unit calculates the material's physical state parameters based on the received physical quantities. This parameter set includes:
[0084] Creep damage measurement parameters ( This parameter is used to quantify the cumulative internal energy dissipation caused by creep deformation in materials; the intermediate health state calculation unit outputs the stress from the stress-strain decoupling unit. ) and creep strain rate ( Using this as input, the real-time rate of change of the creep damage metric parameter is calculated according to the following formula:
[0085] ;
[0086] In the formula, For time; In time Real-time rate of change of creep damage measurement parameters; In time Stress inside the joint; In time The creep strain rate; by numerically integrating this rate of change over time, the current cumulative creep damage metric can be obtained.
[0087] The material's microstructure stability parameter characterizes the crystallinity of the PEEK matrix as a semi-crystalline polymer and the stability of its amorphous regions. The mid-term health state calculation unit includes a non-isothermal crystallization kinetic model, which uses the temperature history provided by the sensor data acquisition module 10 as input to calculate the material's crystallinity in real time. In a specific embodiment, the non-isothermal crystallization kinetic model can be constructed by combining the Avrami equation and the Ozawa equation, or the Mo equation can be used to directly describe the relationship between the crystallization rate and the cooling rate. By using the temperature history provided by the sensor data acquisition module 10 as the model input, the relative crystallinity at the current moment is calculated in real time and output as a material microstructure stability parameter.
[0088] The long-term health status calculation unit is responsible for defining and managing the material baseline physical parameter vector. In one specific embodiment, the parameter vector contains creep coefficients used to describe the Findley creep model. ) and time index ( This model describes the creep compliance of the material. Relationship over time:
[0089] ;
[0090] In the formula, for The creep flexibility at any given moment; This is the initial instantaneous compliance. The creep coefficient; For time exponent. These two parameters ( and This is the core component of the material baseline physical parameter vector. Its initial value is calibrated through offline experiments and updated online by the baseline parameter update module 50 during system operation.
[0091] The parameter vector is a multidimensional vector containing core coefficients describing the slow evolution of the material's macroscopic physical properties over time. In a specific embodiment, the parameter vector includes at least the creep coefficient and time exponent for describing the material's creep compliance. To more comprehensively describe the material's long-term health status, the parameter vector may also include: initial elastic modulus, glass transition temperature, and coefficients related to material fatigue damage. The initial value of the parameter vector is calibrated through offline experiments and updated online by the baseline parameter update module 50 during system operation.
[0092] See attached document Figure 1 and attached Figure 4 Step S3 is by Figure 1 The reinforcement learning decision module 30 shown is executed; the reinforcement learning decision module 30 contains a trained reinforcement learning agent, such as a deep deterministic policy gradient (DDPG) agent or a proximal policy optimization (PPO) agent. This module receives data from the sensor data acquisition module 10 and the object observation module 20, and outputs cooperative control actions to the cooperative control command generation and execution module 40.
[0093] In a specific embodiment employing the DDPG algorithm, the reinforcement learning agent includes an Actor-Network and a Critic-Network. The Actor-Network can be a fully connected neural network with three hidden layers, whose input is a state space vector and whose output is a cooperative control action vector. The Critic-Network's input is the state space vector and the action vector, and its output is the Q-value of the state-action pair. The activation function of all hidden layers can be the Modified Linear Unit (ReLU) function.
[0094] In one specific embodiment, the reinforcement learning agent makes decisions based on an enhanced state space, which is constructed as a vector containing the following three types of data components:
[0095] The first category is real-time sensing data from the sensing data acquisition module 10, specifically including the current joint angle and angular velocity of the robot's carbon fiber joint, as well as strain and temperature values acquired from the distributed fiber optic sensor array.
[0096] The second category consists of material physical state parameters calculated by the material state observation module 20, specifically including the aforementioned creep damage measurement parameters and material microstructure stability parameters.
[0097] The third category is the current material baseline physical parameter vector from the matter observation module 20;
[0098] By integrating these three types of data, the state space provides reinforcement learning agents with comprehensive information about their external motion state and internal material health.
[0099] The cooperative control action output by the reinforcement learning agent is an action vector containing two components. The first component is the torque adjustment action, which is defined as a scalar value used to generate an additional compensating torque. The second component is the physical stiffness adjustment action, which is defined as a scalar value used to set the target stiffness value of the controllable stiffness unit within the joint.
[0100] The training objective of a reinforcement learning agent is to maximize the long-term cumulative value of a compound reward function; refer to Figure 4 In one specific embodiment, the composite reward function ( At each time step The calculation method is as follows:
[0101] ;
[0102] In the formula: For time; In time Total reward value; In time The task performance reward item is used to characterize the tracking accuracy of the joint motion trajectory; in one embodiment, the item is calculated as the negative square of the error between the expected joint trajectory and the actual joint trajectory. In time A mid-term health maintenance bonus is provided, which is used to penalize the deterioration of the material's physical state parameters; in one embodiment, the bonus is calculated as a value of the creep damage metric parameter or a negative value of its rate of change. In time A long-term health gain reward term is provided, which is used to penalize excessively rapid changes in the material baseline physical parameter vector; in one embodiment, this term is calculated as the negative of the time rate of change norm of the material baseline physical parameter vector. , ,and The preset non-negative weighting coefficient is used to balance the relative importance of the three reward items, and the sum of the three is 1.
[0103] See attached document Figure 1 and attached Figure 5 Step S4 is by Figure 1 The cooperative control command generation and execution module 40 shown is executed; this module receives the cooperative control action output by the reinforcement learning decision module 30, and parses it into two independent control commands: a command torque and a command stiffness, and then drives the actuator of the robot's carbon fiber joint to execute synchronously.
[0104] See attached document Figure 5 In one specific embodiment, the process of generating the command torque for driving the joint includes the following operations:
[0105] First, the collaborative control command generation and execution module 40 calculates a predicted creep displacement based on the current material baseline physical parameter vector and stress and temperature state obtained from the physical state observation module 20 through a built-in short-term creep prediction model. The predicted creep displacement is superimposed on the original expected trajectory given by the upper-level motion planner to obtain a corrected expected trajectory. This is a feedforward compensation operation.
[0106] Secondly, the cooperative control command generation and execution module 40 provides the corrected desired trajectory as input to a base controller, which can be a proportional-integral-derivative (PID) controller or a linear quadratic regulator (LQR). The base controller calculates and generates a base torque based on the deviation between the corrected desired trajectory and the actual joint position.
[0107] See attached document Figure 1 and attached Figure 5 Step S4 is by Figure 1The cooperative control instruction generation and execution module 40 shown executes the cooperative control instruction generation and execution module 40; the cooperative control instruction generation and execution module 40 receives the cooperative control action from the reinforcement learning decision module 30, and generates physically executable instruction torque and instruction stiffness accordingly, and finally sends these two instructions synchronously to the drive motor and controllable stiffness unit of the robot joint.
[0108] In one specific embodiment, the process of generating the command torque combines model-based feedforward compensation and reinforcement learning-based feedback regulation. (See Appendix) Figure 5 The process specifically includes:
[0109] First, the collaborative control command generation and execution module 40 uses the creep strain calculated by the physical state observation module 20 ( The predicted creep displacement is calculated using the following error propagation model. ):
[0110] ;
[0111] In the formula, This represents the predicted joint angle offset. These are coefficients related to joint geometry; In time creep strain; In time of, and temperature The relevant material elastic modulus.
[0112] In calculating the predicted creep displacement ( After that, it is superimposed on the original expected trajectory given by the robot task planning layer. On ), thus obtaining a corrected expected trajectory ( The calculation method is as follows:
[0113] ;
[0114] The corrected desired trajectory is fed into a base controller, such as a proportional-integral-derivative (PID) controller. This base controller calculates a base torque based on the error between the corrected desired trajectory and the trajectory fed back by the joint's actual position sensor. This torque is used to complete the main trajectory tracking task. Simultaneously, the cooperative control command generation and execution module 40 converts the torque adjustment component in the cooperative control action output by the reinforcement learning decision module 30 into a compensation torque with physical units. ).
[0115] As a specific method of torque compensation, this compensation torque ( ) can be based on the predicted creep displacement ( ,Right now ), generated by a proportional-derivative (PD) controller:
[0116] ;
[0117] In the formula, To compensate for torque; This represents the predicted creep displacement. Its rate of change; and These are the preset proportional and differential gain coefficients.
[0118] The base torque output from the base controller is algebraically summed with the compensation torque to obtain the final command torque. This command torque is then sent to the drive motor of the joint for execution.
[0119] The process of generating command stiffness involves mapping the physical stiffness adjustment component in the cooperative control action output by the reinforcement learning decision module 30 to the actual physical stiffness value of the controllable stiffness unit. In a specific embodiment, the physical stiffness adjustment action is a value normalized to a range. The cooperative control command generation and execution module 40 sets an achievable minimum stiffness value based on the physical characteristics of the controllable stiffness unit (e.g., a magnetorheological damper or a piezoelectric stack actuator) within the joint. ) and maximum stiffness value ( ).
[0120] The collaborative control instruction generation and execution module 40 uses the following linear mapping relationship to generate and execute physical stiffness adjustment values ( ) is converted into the final command stiffness ( ):
[0121] ;
[0122] In the formula: The final generated command stiffness; This represents the minimum physical stiffness value achievable by a controllable stiffness element. This represents the maximum physical stiffness value achievable by a controllable stiffness element. The physical stiffness adjustment action value received from the reinforcement learning decision module 30, within the range of values.
[0123] In one specific embodiment, when the controllable stiffness unit is a magnetorheological damper, the cooperative control command generation and execution module 40 further includes a current driving unit, which calculates the command stiffness ( Based on the pre-calibrated stiffness and current characteristic curves, a corresponding driving current value is converted and applied to the electromagnetic coil of the magnetorheological damper.
[0124] The calculated command stiffness value is sent to the controller of the controllable stiffness unit to adjust its physical stiffness.
[0125] See attached document Figure 1 Step S5 is by Figure 1 The baseline parameter update module 50 shown is executed; the function of the baseline parameter update module 50 is to update the material baseline physical parameter vector stored in the physical state observation module 20 online and in a closed loop according to the actual load history of the robot carbon fiber joint during long-term service, so as to reflect the evolution of the physical properties of the material due to aging, fatigue and other factors.
[0126] The baseline parameter update module 50 is connected to the collaborative control command generation and execution module 40 to receive the execution history data of command torque and command stiffness. This module contains a data storage unit for continuously recording the stress history and temperature history of the joint. The stress history is calculated by combining the received command torque history with the known joint geometry model; the temperature history comes directly from the sensor data acquisition module 10.
[0127] In one specific embodiment, the parameter update is performed periodically; the baseline parameter update module 50 sets a preset time window, for example, every 100 working hours. Within this time window, the module continuously accumulates the stress and temperature data of the joint, and when a time window ends, an update calculation is triggered.
[0128] The core of the update calculation is based on an evolution model of a material baseline physical parameter vector. This model describes the relationship between the rate of change of the parameter vector and its experienced load history and current state. In one embodiment, this evolution model is expressed as:
[0129] ;
[0130] In the formula: For time; In time The material baseline physical parameter vector; In time The rate of change of the material baseline physical parameter vector; The preset evolution function is a vector function whose specific form is constructed based on prior knowledge of materials science. For example, for the temperature-related components in the parameter vector, the evolution can be described by the Arrhenius model to describe the effect of temperature-accelerated aging; for the stress-related components, the evolution can be described by a model based on damage mechanics or fatigue accumulation. Deadline The joint stress history is the stress time series recorded by the baseline parameter update module 50; Deadline The joint temperature history is the temperature time series recorded by the baseline parameter update module 50.
[0131] For example, for the material baseline physical parameter vector creep coefficient Components, their evolution function A specific form can be expressed as:
[0132] ;
[0133] In the formula, The rate of change of the creep coefficient; These are the material constants determined experimentally. and These are the average stress and average absolute temperature in the previous update cycle, respectively. This is the aging activation energy of the material; The constant is the ideal gas constant. Using this model, the evolution increment of the creep coefficient over time can be quantitatively calculated based on recorded stress and temperature history.
[0134] Each time an update is triggered, the baseline parameter update module 50 takes the stress history and temperature history recorded within the time window, as well as the current material baseline physical parameter vector, as input, substitutes them into the above evolution model formula, and calculates the increment of the parameter vector within the time window using a numerical integration method (e.g., the first-order Euler method or the fourth-order Runge-Kutta method). Then, it adds the increment to the old parameter vector to obtain the updated new parameter vector.
[0135] After the calculation is completed, the baseline parameter update module 50 writes the updated material baseline physical parameter vector into the designated storage area of the state observation module 20 for use in the state observation calculation in the next working cycle.
[0136] See attached document Figure 1 - Appendix Figure 5 To more comprehensively illustrate the collaborative working process of the technical solution of this invention, a specific application scenario will be used as an example below. The scenario is set as follows: a six-axis robot equipped with the control system of this invention performs a cyclical high-load handling task at its wrist carbon fiber joint. The task involves continuously moving a heavy object back and forth between point A and point B.
[0137] In the initial stage of task execution, the robot joints are in a completely new state. When the robot begins its first handling cycle, the system operates as follows:
[0138] When the joint begins to bear load and move, the sensor data acquisition module 10 collects strain and temperature data of each measuring point inside the joint in real time. In the initial state, the temperature is the ambient temperature and the total strain is mainly elastic strain.
[0139] The physical state observation module 20 receives this data, and its internal stress-strain decoupling unit calculates the stress distribution inside the joint. The creep damage measurement parameter calculated by the mid-term health state calculation unit has an initial value of zero, and the material baseline physical parameter vector in the long-term health state calculation unit is the initial calibration value.
[0140] The reinforcement learning decision module 30 takes the state vector, which includes the current joint angle, angular velocity, zero-damage metric, and initial baseline parameters, as input. Since the material health state parameters are all at optimal values, the penalty values of the mid-term health maintenance reward and the long-term health gain reward in the composite reward function are both zero. At this time, the task performance reward dominates. Therefore, the cooperative control action output by the reinforcement learning agent will mainly aim to maximize the trajectory tracking accuracy. For example, it will output a torque adjustment action to generate a large acceleration torque and a physical stiffness adjustment action set to a moderate value.
[0141] Based on this action, the collaborative control command generation and execution module 40 generates a command torque sufficient to drive the joint to move rapidly and a command stiffness of moderate magnitude, and executes them synchronously.
[0142] After the robot has worked continuously for several hours, the joint temperature rises due to heat generated by the motors and the movement of molecular chains within the material. Simultaneously, continuous operation under high stress causes material creep. At this point, the system's workflow exhibits adaptive adjustment:
[0143] The sensor data acquisition module 10 detects a significant increase in the internal temperature of the joint, and can still detect minute residual strain when the joint is unloaded (e.g., when a heavy object is placed at point A or point B). After receiving new sensor data, the physical state observation module 20, through its intermediate health state calculation unit 22, calculates a non-zero and continuously increasing creep damage metric parameter based on the stress history and the integral of the creep strain rate.
[0144] The reinforcement learning decision module 30 receives a state vector containing an increased temperature and a non-zero creep damage metric parameter. At this point, the mid-term health maintenance reward term in the composite reward function begins to generate a negative reward value. In order to maximize the long-term cumulative reward, the agent's policy will be adjusted to output a new cooperative control action. For example, this action may correspond to a slight negative compensation torque to smooth the motion curve and reduce peak stress; at the same time, it may correspond to a larger physical stiffness adjustment action value to enhance the physical support of the joint and resist creep deformation.
[0145] Based on this new action, the collaborative control command generation and execution module 40 generates a command torque with a slightly lower peak value and simultaneously generates a higher command stiffness, which is sent to the controllable stiffness unit such as the magnetorheological damper for execution. In this way, the system actively slows down the rate of material damage accumulation while sacrificing a very small amount of task execution speed.
[0146] When the system has accumulated a preset long time (e.g., 1000 hours), the baseline parameter update module 50 is triggered to perform an online update.
[0147] This module reads the complete stress and temperature history data stored internally over the past 1000 hours; then, it substitutes this historical data, along with the current material baseline physical parameter vector, into the preset evolution model formula of the material baseline physical parameter vector for calculation. The calculation results reflect the changes in the physical properties of the material after long-term service, such as a small increase in the creep coefficient or a small decrease in the elastic modulus.
[0148] The baseline parameter update module 50 writes the calculated new material baseline physical parameter vector into the storage area of the physical state observation module 20.
[0149] After this update, the entire system continues to perform tasks, but its internal model has adapted to the aging of the material. In subsequent calculations, the material state observation module 20 will use this updated parameter vector, which is closer to the actual physical state of the material, to more accurately decouple strain and predict creep. This makes the decision basis of the reinforcement learning decision module 30 more accurate, and the feedforward compensation of the cooperative control command generation and execution module 40 more precise, thereby ensuring that the robot can maintain optimal task performance and joint health balance at different stages throughout its life cycle.
[0150] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for coordinated dynamic load control of carbon fiber joints in robots based on reinforcement learning, characterized in that, The method includes the following steps: S1. Acquire sensing data within the robot's carbon fiber joint, the sensing data including strain and temperature; S2. Based on the sensing data, calculate the material physical state parameters of the joint material in the mid-term health state and the material baseline physical parameter vector of the joint material in the long-term health state through the physical state observation model. S3. The sensing data, material physical state parameters, and material baseline physical parameter vector are used together as the state input of the reinforcement learning agent, and the reinforcement learning agent outputs a coordinated control action including torque adjustment action and physical stiffness adjustment action. S4. Based on the coordinated control action, generate the command torque to drive the joint and the command stiffness to adjust the controllable stiffness unit within the joint, and simultaneously execute the command torque and command stiffness on the joint. S5. Based on the execution history of the command torque and command stiffness, update the material baseline physical parameter vector in the physical state observation model, and provide the updated material baseline physical parameter vector to the physical state observation model for subsequent calculations; In step S2, the material physical state parameters of the joint material in the mid-term health state are calculated by the physical state observation model, including: creep damage measurement parameters, which are used to quantify the cumulative energy dissipation caused by creep; and material microstructure stability parameters, which are used to characterize the crystallinity and stability of the amorphous region of the material. In step S2, the material baseline physical parameter vector of the joint material in the long-term health state includes the creep coefficient, time exponent and initial elastic modulus used to describe the material creep compliance.
2. The method for coordinated dynamic load control of robot carbon fiber joints based on reinforcement learning according to claim 1, characterized in that, Step S1, the step of acquiring the sensing data within the robot's carbon fiber joint, includes: The strain and temperature are collected by a distributed fiber optic grating sensor array embedded inside the robot's carbon fiber joint structure.
3. The method for coordinated dynamic load control of robot carbon fiber joints based on reinforcement learning according to claim 1, characterized in that, The calculation steps for the creep damage measurement parameters include: Based on the sensor data, the stress and creep strain rate inside the joint are calculated. Based on the stress and creep strain rate, the creep damage measurement parameter is calculated, and the real-time change rate of the creep damage measurement parameter satisfies the following formula: ; In the formula, For time; The real-time rate of change of the creep damage measurement parameter; This refers to the stress inside the joint; denoted as creep strain rate.
4. The method for coordinated dynamic load control of robot carbon fiber joints based on reinforcement learning according to claim 1, characterized in that, Step S3 specifically includes: The reinforcement learning agent is trained by maximizing the long-term cumulative value of a composite reward function, which includes: A performance reward item used to characterize the accuracy of joint motion trajectory tracking; Mid-term health maintenance bonus item, which is negatively correlated with the degree of degradation of the material's physical state parameters; The long-term health gain reward term is negatively correlated with the time rate of change of the material baseline physical parameter vector.
5. The method for coordinated dynamic load control of robot carbon fiber joints based on reinforcement learning according to claim 1, characterized in that, The controllable stiffness unit within the joint mentioned in step S4 is a magnetorheological damper, an electrorheological damper, or a piezoelectric stack actuator.
6. The method for coordinated dynamic load control of robot carbon fiber joints based on reinforcement learning according to claim 1, characterized in that, Step S4, the step of generating the command torque to drive the joint, includes: The predicted creep displacement is superimposed on the desired trajectory to obtain the corrected desired trajectory; The proportional-integral-differential (PI-DI) fundamental controller generates the fundamental torque based on the corrected desired trajectory. The command torque is obtained by adding the base torque to the compensation torque generated according to the torque adjustment action.
7. The method for coordinated dynamic load control of robot carbon fiber joints based on reinforcement learning according to claim 1, characterized in that, Step S5, updating the material baseline physical parameter vector within the physical state observation model, includes: The update of the material baseline physical parameter vector is based on the following evolution model: ; In the formula, For time; The rate of change of the material baseline physical parameter vector; This is a preset evolution function; For the stress history of the joint; The temperature history of the joint; In time The material baseline physical parameter vector; The evolution function is based on the Arrhenius model and is used to characterize the accelerating effect of temperature on material aging.
8. A reinforcement learning-based dynamic load cooperative control system for robot carbon fiber joints, applied to the method described in any one of claims 1-7, characterized in that, The system includes: The sensor data acquisition module is used to acquire sensor data within the robot's carbon fiber joint, including strain and temperature. The physical state observation module is used to receive the sensing data and calculate the material physical state parameters of the joint material in the mid-term health state and the material baseline physical parameter vector of the joint material in the long-term health state based on the sensing data. The reinforcement learning decision module is used to receive the sensing data, material physical state parameters, and material baseline physical parameter vector as state inputs, and output a coordinated control action including torque adjustment action and physical stiffness adjustment action. The collaborative control command generation and execution module is used to generate and synchronously execute the command torque for driving the joint and the command stiffness for adjusting the controllable stiffness unit within the joint, based on the collaborative control action. The baseline parameter update module is used to update the material baseline physical parameter vector in the physical state observation module according to the execution history of the command torque and command stiffness, and provide the updated material baseline physical parameter vector to the physical state observation module.
Citation Information
Patent Citations
Robot state prediction method and device, electronic equipment and storage medium
CN117733874A
Industrial robot control system and method
CN120255424A