Robot carbon fiber joint dynamic load cooperative control method based on reinforcement learning

By employing a reinforcement learning-based dynamic load collaborative control method for robot carbon fiber joints, material creep and aging are monitored and compensated in real time. This solves the problems of decreased control accuracy and reduced reliability in existing technologies, and achieves long-term health maintenance and consistent control performance of robot joints.

CN120816501AActive Publication Date: 2025-10-21JILIN UNIVERSITY

Patent Information

Application Number
CN202511324363.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-10-21
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing technologies for robot control of carbon fiber composite joints lack effective prediction and compensation for time-varying effects such as material creep and aging, leading to a long-term decrease in control accuracy and reliability.

Method used

A reinforcement learning-based dynamic load collaborative control method for carbon fiber robot joints is adopted. By acquiring strain and temperature within the joint through sensor data, calculating material health parameters using a physical observation model, and combining reinforcement learning agents to generate collaborative control actions for torque and physical stiffness adjustment, the baseline parameters of the material are updated in real time, thereby achieving proactive compensation for material performance degradation.

Benefits of technology

It improves the control robustness under conditions of changing material properties, ensures the consistency and long-term effectiveness of control performance throughout the robot joint's life cycle, and extends its service life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120816501A_ABST
    Figure CN120816501A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of humanoid robots, and discloses a robot carbon fiber joint dynamic load cooperative control method based on reinforcement learning, and the method comprises the steps: obtaining joint strain and temperature; calculating medium-term state parameters and long-term baseline parameters of the material; the sensing data and the parameters are used as state input, and torque and rigidity are output to act cooperatively; generating and synchronously executing instruction torque and instruction rigidity; and updating the baseline parameters on line according to the execution history to form closed-loop adaptive control. The system comprises a sensing data acquisition module, a physical state observation module, a reinforcement learning decision module, a cooperative control instruction generation and execution module and a baseline parameter updating module. According to the invention, through cooperative control of torque and rigidity and closed-loop updating of a material state, material performance evolution can be actively compensated, and the service life of the joint is prolonged while the task precision is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of humanoid robots, and in particular to a method for collaboratively controlling dynamic loads of robot carbon fiber joints based on reinforcement learning. Background Art

[0002] In recent years, with the development of robotics technology towards high dynamics, high load and lightweight, carbon fiber composite materials have been increasingly widely used in robot structural design due to their excellent properties of high specific strength and high specific modulus, especially in key load-bearing parts such as the limbs and joints of humanoid robots.

[0003] Existing control models for robots using carbon fiber composite joints are often based on the assumption of ideal rigid-body dynamics, simplifying the robot joints into components with constant physical properties. This control approach can achieve good trajectory tracking performance during short-term, low-load tasks. However, carbon fiber reinforced polymer-based composites have inherent viscoelastic properties, and when subjected to long-term cyclic loads or sustained high-temperature operating environments, they can exhibit time-varying effects such as creep, stress relaxation, and material aging.

[0004] Existing control strategies typically prioritize instantaneous task execution accuracy, but lack effective online observation and feedforward compensation mechanisms for slowly accumulating end-effector position deviations caused by factors such as material creep. Furthermore, control systems often fail to fully consider the evolution of the internal health status of joint materials when making decisions. The generation of control laws typically does not include the goal of proactively maintaining the long-term health of materials and slowing performance degradation. This can lead to problems such as decreased control accuracy, permanent joint deformation, and reduced reliability after long-term service.

[0005] Therefore, the present invention proposes a robot carbon fiber joint dynamic load collaborative control method based on reinforcement learning to address the shortcomings of the existing technology. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention provides a dynamic load collaborative control method for robot carbon fiber joints based on reinforcement learning. It solves the problem that while the robot control method pursues instantaneous task accuracy, it lacks effective prediction and collaborative compensation means for the performance degradation of carbon fiber joints caused by time-varying effects such as creep and aging, making it difficult to balance the short-term control performance and long-term structural health and reliability of the robot.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: a robot carbon fiber joint dynamic load collaborative control method based on reinforcement learning, the method comprising the following steps: S1. Acquire sensor data in the carbon fiber joint of the robot, wherein the sensor data includes strain and temperature; S2. Calculating, based on the sensor data, material physical state parameters of the joint material in a mid-term health state and a material baseline physical parameter vector of the joint material in a long-term health state using a physical state observation model; S3. Using the sensor data, the material physical state parameters, and the material baseline physical parameter vector as state inputs of a reinforcement learning agent, and having the reinforcement learning agent output a coordinated control action including a torque adjustment action and a physical stiffness adjustment action; S4. generating a command torque for driving the joint and a command stiffness for adjusting a controllable stiffness unit in the joint according to the coordinated control action, and synchronously executing the command torque and command stiffness on the joint; S5. Update the material baseline physical parameter vector in the physical state observation model according to the execution history of the command torque and the command stiffness, and provide the updated material baseline physical parameter vector to the physical state observation model for subsequent calculations.

[0008] Preferably, in step S1, the step of obtaining sensor data in the carbon fiber joint of the robot includes: The strain and temperature are collected by a distributed fiber grating sensor array embedded in the carbon fiber joint structure of the robot.

[0009] Preferably, in step S2, calculating and obtaining the material physical state parameters of the mid-term health state of the joint material through the physical state observation model includes: Creep damage metric, used to quantify the cumulative energy dissipation of the material due to creep; The material microstructure stability parameter is used to characterize the stability of the material's crystallinity and amorphous regions.

[0010] Preferably, the step of calculating the creep damage metric parameter includes: Calculating the stress and creep strain rate inside the joint based on the sensing data; The creep damage metric parameter is calculated based on the stress and creep strain rate. The real-time change rate of the creep damage metric parameter satisfies the following formula: ; Where, For time; is the real-time change rate of creep damage measurement parameters; is the stress inside the joint; is the creep strain rate.

[0011] Preferably, in step S2, the material baseline physical parameter vector of the long-term health state of the joint material at least includes a creep coefficient, a time index and an initial elastic modulus for describing the creep compliance of the material.

[0012] Preferably, step S3 specifically includes: The reinforcement learning agent is trained by maximizing the long-term cumulative value of a compound reward function, which includes: Task performance reward item used to characterize the accuracy of joint motion trajectory tracking; A mid-term health maintenance bonus item, wherein the mid-term health maintenance bonus item is negatively correlated with the degree of deterioration of the material's physical state parameters; A long-term health gain reward item, wherein the long-term health gain reward item is negatively correlated with the time rate of change of the material baseline physical parameter vector.

[0013] Preferably, the controllable stiffness unit in the joint in step S4 is a magnetorheological damper, an electrorheological damper or a piezoelectric stack actuator.

[0014] Preferably, in step S4, the step of generating a command torque for driving the joint includes: The predicted creep displacement is superimposed on the expected trajectory to obtain a corrected expected trajectory; generating a base torque according to the modified desired trajectory by a proportional-integral-derivative base controller; The basic torque is added to the compensation torque generated according to the torque adjustment action to obtain the command torque.

[0015] Preferably, in step S5, the step of updating the material baseline physical parameter vector in the physical state observation model includes: The update of the material baseline physical parameter vector is based on the following evolution model: ; Where, For time; is the rate of change of the material baseline physical parameter vector; is the preset evolution function; is the stress history of the joint; Temperature history for joints; For in time The material baseline physical parameter vector; The evolution function is a function based on the Arrhenius model, which is used to characterize the accelerated effect of temperature on material aging.

[0016] The present invention also provides a robot carbon fiber joint dynamic load collaborative control system based on reinforcement learning, the system comprising: A sensor data acquisition module is used to obtain sensor data in the carbon fiber joints of the robot, wherein the sensor data includes strain and temperature; a physical state observation module, configured to receive the sensor data and calculate, based on the sensor data, material physical state parameters of the joint material in the mid-term health state and a material baseline physical parameter vector of the joint material in the long-term health state; a reinforcement learning decision module, configured to receive the sensor data, the material physical state parameters, and the material baseline physical parameter vector as state inputs, and output a coordinated control action including a torque adjustment action and a physical stiffness adjustment action; a collaborative control instruction generation and execution module, configured to generate and synchronously execute an instruction torque for driving the joint and an instruction stiffness for adjusting a controllable stiffness unit in the joint according to the collaborative control action; The baseline parameter updating module is used to update the material baseline physical parameter vector in the physical state observation module according to the execution history of the command torque and the command stiffness, and provide the updated material baseline physical parameter vector to the physical state observation module.

[0017] The present invention provides a method for collaborative dynamic load control of robot carbon fiber joints based on reinforcement learning. It has the following beneficial effects: 1. The present invention constructs a physical state observation model to calculate in real time the material physical state parameters that characterize the mid-term health state of the joint material and the material baseline physical parameter vectors that characterize the long-term health state, and uses these parameters as the input of the reinforcement learning agent. This setting enables the control decision to distinguish the source of task execution errors, such as whether it is caused by improper control instructions or material creep, thereby performing more targeted control adjustments and improving the control robustness under conditions of changing material properties.

[0018] 2. The present invention generates and synchronously executes command torque and command stiffness by outputting collaborative control actions including torque adjustment actions and physical stiffness adjustment actions through reinforcement learning intelligent body. This collaborative control method gives the system multi-dimensional adjustment capabilities for handling dynamic loads, enabling it to collaboratively optimize active torque compensation and passive stiffness adjustment according to different task requirements and joint health status, thereby reducing stress impact on the joint structure while ensuring task execution.

[0019] 3. The present invention updates the material baseline physical parameter vector in the physical state observation model online based on the execution history of the command torque and command stiffness. This mechanism establishes a closed-loop control model that adapts to the actual evolution of the material's physical properties, enabling the control system to track and adapt to irreversible changes in the material due to factors such as aging and fatigue, ensuring the consistency and long-term effectiveness of control performance throughout the life cycle of the robot joint.

[0020] 4. The present invention introduces penalty terms for the medium-term and long-term health status of the material into the composite reward function of reinforcement learning, guiding the optimization direction of the control strategy from a single task accuracy target to the comprehensive optimum between task accuracy and long-term joint health. This enables the system to learn and execute material-friendly load modes, actively avoid or slow down operations that cause material performance degradation, and thus effectively extend the service life of the robot's carbon fiber joints. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a structural block diagram of the robot carbon fiber joint dynamic load collaborative control system of the present invention; Figure 2 This is a flow chart of the method for collaboratively controlling dynamic loads of carbon fiber joints of a robot according to the present invention; Figure 3 This is a block diagram of the internal structure of a specific embodiment of the physical state observation module of the present invention; Figure 4 Schematic diagram of the structure of the composite reward function of the present invention; Figure 5 Schematic diagram of the command torque generation process of the present invention.

[0022] Among them, 10, sensor data acquisition module; 20, physical state observation module; 30, reinforcement learning decision module; 40, collaborative control instruction generation and execution module; 50, baseline parameter update module. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0024] Refer to the attached Figure 1 , Figure 1 1 is a structural block diagram of a robot carbon fiber joint dynamic load collaborative control system according to an embodiment of the present invention; the present invention provides a robot carbon fiber joint dynamic load collaborative control system, which may include: a sensor data acquisition module 10, a physical state observation module 20, a reinforcement learning decision module 30, a collaborative control instruction generation and execution module 40, and a baseline parameter update module 50; The sensor data acquisition module 10 is configured to acquire sensor data within the carbon fiber joints of the robot, wherein the sensor data includes strain and temperature. The sensor data acquisition module 10 transmits the acquired sensor data to the physical state observation module 20 and the reinforcement learning decision module 30 respectively. The physical state observation module 20 is connected to the sensor data acquisition module 10 and is used to receive sensor data. The physical state observation module 20 is configured to calculate and obtain material physical state parameters representing the mid-term health state of the joint material and a material baseline physical parameter vector representing the long-term health state of the joint material based on the sensor data. The calculated results are transmitted to the reinforcement learning decision module 30. The reinforcement learning decision module 30 is connected to the sensor data acquisition module 10 and the physical state observation module 20 respectively. The reinforcement learning decision module 30 is configured to use the received sensor data, material physical state parameters, and material baseline physical parameter vector as state inputs, and output a coordinated control action including a torque adjustment action and a physical stiffness adjustment action based on the state inputs. The coordinated control action is transmitted to the coordinated control instruction generation and execution module 40. The collaborative control instruction generation and execution module 40 is connected to the reinforcement learning decision module 30. The collaborative control instruction generation and execution module 40 is configured to generate a command torque for driving the joint and a command stiffness for adjusting the controllable stiffness unit in the joint based on the received collaborative control action, and synchronously execute the command torque and command stiffness on the joint. At the same time, the collaborative control instruction generation and execution module 40 transmits the execution history data of the command torque and command stiffness generated by it to the baseline parameter update module 50. The baseline parameter update module 50 is connected to the collaborative control instruction generation and execution module 40; the baseline parameter update module 50 is configured to update the material baseline physical parameter vector stored inside the physical state observation module 20 according to the execution history of the received instruction torque and instruction stiffness; through this update, the physical state observation module 20 will use the updated parameters in subsequent calculations, thereby forming a closed-loop regulation path.

[0025] Refer to the attached Figure 2 , Figure 2 1 is a flow chart of a method for collaboratively controlling dynamic loads of a robot carbon fiber joint according to an embodiment of the present invention. The method generally comprises the following steps: Step S1, acquiring sensor data in the carbon fiber joint of the robot, the sensor data including strain and temperature; Step S2, inputting the sensor data into a physical state observation model, and calculating the physical state parameters of the joint material in the mid-term health state and the material baseline physical parameter vector of the joint material in the long-term health state by the physical state observation model; Step S3: The sensor data, the material physical state parameters, and the material baseline physical parameter vector are used as the state input of the reinforcement learning agent, and the reinforcement learning agent outputs a coordinated control action including a torque adjustment action and a physical stiffness adjustment action; Step S4, generating a command torque for driving the joint and a command stiffness for adjusting the controllable stiffness unit in the joint according to the coordinated control action, and synchronously executing the command torque and command stiffness on the joint; Step S5: updating the material baseline physical parameter vector in the physical state observation model according to the execution history of the command torque and the command stiffness, so as to provide the physical state observation model with use in subsequent calculations.

[0026] The following will describe in detail the various components and steps of the robot carbon fiber joint dynamic load collaborative control system and method provided by the present invention.

[0027] Refer to the attached Figure 1 , step S1 is composed of Figure 1 The sensor data acquisition module 10 shown is executed; in a specific embodiment, the sensor data acquisition module 10 includes a distributed fiber Bragg grating sensor array, a broadband light source, and a fiber Bragg grating demodulator; the distributed fiber Bragg grating sensor array is directly embedded into the interlayer structure of the carbon fiber reinforced polymer matrix during the manufacturing process of the robot carbon fiber joint; the array is composed of multiple optical fibers, each of which is inscribed with multiple fiber Bragg gratings (FBGs) with different reflection wavelengths. These gratings are distributed along a preset path to cover key stress concentration areas and temperature sensitive areas within the joint; When the system is in operation, a broadband light source emits a beam of broad-spectrum light, which is injected into the optical fiber of the distributed fiber Bragg grating sensor array. As the light beam propagates through the optical fiber, each fiber Bragg grating reflects light of its specific central wavelength (i.e., the Bragg wavelength) back into the optical path. Changes in the strain or temperature of the joint material cause the period or effective refractive index of the grating to change, resulting in a shift in the Bragg wavelength of its reflection. The function of the fiber Bragg grating interrogator is to detect the spectrum reflected from each fiber Bragg grating in real time with high precision and accurately calculate the Bragg wavelength drift of each grating. Because the wavelength drift of a single fiber Bragg grating is affected by both strain and temperature, this embodiment uses a temperature compensation configuration to accurately separate these two physical quantities. In one specific configuration, a compensation grating for temperature measurement is placed near each strain-measuring grating. This compensation grating is encapsulated in a micro-capillary tube that is immune to mechanical stress, ensuring that it responds only to temperature changes. Therefore, the fiber Bragg grating interrogator first calculates the precise temperature at the compensation grating's location based on the wavelength shift of the compensation grating. The fiber Bragg grating interrogator then uses the known temperature-wavelength drift relationship to subtract the drift component caused by temperature changes from the total wavelength drift of the paired strain measurement grating. After this compensation calculation, the remaining wavelength drift is the drift caused solely by mechanical strain. The precise strain value at that point can be calculated using the calibrated strain-wavelength drift relationship. Finally, the sensor data acquisition module 10 integrates the data demodulated from all grating points to generate a discrete strain field distribution map and temperature field distribution map that represent the internal state of the entire joint; these two sets of sensor data containing spatial position information will be output as digital signals to the physical state observation module 20 and the reinforcement learning decision module 30.

[0028] In other embodiments, the sensor data acquisition module 10 may also employ other types of sensors to acquire strain and temperature data. For example, a patch-type or sputtered film piezoresistive strain gauge array may be used to measure strain, combined with a thermistor or thermocouple to measure temperature. Furthermore, data transmission between the sensor data acquisition module 10, the physical state observation module 20, and the reinforcement learning decision module 30 may be implemented via a high-speed industrial bus (e.g., CAN bus or EtherCAT) to ensure real-time and reliable data.

[0029] Refer to the attached Figure 1 and attached Figure 3 , step S2 is performed by Figure 1 The physical state observation module 20 shown is executed; the physical state observation module 20 receives the sensor data provided by the sensor data acquisition module 10. Its core function is to convert the original strain and temperature data with relatively simple physical meaning into a multi-dimensional parameter set that can comprehensively characterize the internal health state of the robot carbon fiber joint material, including the material physical state parameters representing the medium-term state and the material baseline physical parameter vector representing the long-term state; In a specific embodiment, the dynamic behavior of the robot carbon fiber joint can be described by the following equation: ; Where, , , are the angle, angular velocity and angular acceleration vector of the joint respectively; is the mass inertia matrix; are the Coriolis force and centrifugal force terms; is the gravity term; is the motor output torque; is the equivalent interference torque generated by the creep effect of the material. Perform precise modeling and compensation.

[0030] Refer to the attached Figure 3 ,In a specific embodiment, the internal functions of the physical state observation module 20 can be divided into: a stress-strain decoupling unit, a mid-term health state calculation unit, and a long-term health state calculation unit; The stress-strain decoupling unit receives the total strain ( ) and temperature ( ); Since the total strain is the elastic strain ( ) and creep strain ( ), that is, , and the subsequent health status assessment needs to distinguish different strain components and their corresponding stresses, so the unit first performs decoupling calculations; the unit reads the current material baseline physical parameter vector updated by the baseline parameter update module 50 from the internal storage, and obtains the real-time elastic modulus ( ); Then, by solving the constitutive relations of the materials (for example, ) and creep models (e.g., Findley creep model), separating the current internal stress from the total strain ( ) and creep strain ( ); The creep strain rate is calculated by numerically differentiating the historical creep strain series ( ); The element will calculate the stress ( ), creep strain rate ( ) and temperature ( ) is transmitted to the mid-term health status calculation unit.

[0031] The mid-term health status calculation unit calculates the material physical state parameters based on the received physical quantities. The parameter set includes: Creep damage metric parameters ( This parameter is used to quantify the cumulative internal energy dissipation of the material due to creep deformation; the medium-term health state calculation unit uses the stress output by the stress-strain decoupling unit ( ) and creep strain rate ( ) as input, the real-time rate of change of creep damage metric parameters is calculated according to the following formula: ; Where, For time; For in time Real-time rate of change of creep damage measurement parameters; For in time stress within the joint; For in time The creep strain rate is the creep strain rate; by numerically integrating the change rate over time, the current cumulative creep damage metric can be obtained.

[0032] The material microstructure stability parameter is used to characterize the crystallinity of the PEEK matrix as a semi-crystalline polymer and the stability of the amorphous region. The mid-term health status calculation unit contains a non-isothermal crystallization kinetic model. The model uses the temperature history provided by the sensor data acquisition module 10 as input to calculate the crystallinity of the material in real time. ) evolution process, in a specific embodiment, the non-isothermal crystallization kinetics model can be constructed by combining the Avrami equation and the Ozawa equation, or the Mo equation can be used to directly describe the relationship between the crystallization rate and the cooling rate; by using the temperature history provided by the sensor data acquisition module 10 as the model input, the relative crystallinity at the current moment is calculated by real-time integration and output as a material microstructure stability parameter.

[0033] The long-term health calculation unit is responsible for defining and managing the material baseline physical parameter vector ( In a specific embodiment, the parameter vector includes creep coefficients for describing the Findley creep model ( ) and the time index ( ). This model describes the creep compliance of the material Relationship over time: ; Where, for Creep compliance at time is the initial instantaneous compliance; is the creep coefficient; is the time index. These two parameters ( and ) is the core component of the material baseline physical parameter vector, whose initial value is calibrated through offline experiments and is updated online by the baseline parameter update module 50 during system operation.

[0034] The parameter vector is a multidimensional vector that contains core coefficients that describe the slow evolution of the macroscopic physical properties of the material over time. In a specific embodiment, the parameter vector includes at least a creep coefficient and a time index for describing the creep compliance of the material. To more comprehensively describe the long-term health of the material, the parameter vector may also include: an initial elastic modulus, a glass transition temperature, and coefficients related to fatigue damage of the material. The initial value of the parameter vector is calibrated through offline experiments and is updated online by the baseline parameter update module 50 during system operation.

[0035] Refer to the attached Figure 1 and attached Figure 4 , step S3 is performed by Figure 1The reinforcement learning decision module 30 shown is executed; the reinforcement learning decision module 30 contains a trained reinforcement learning agent, such as a deep deterministic policy gradient (DDPG) agent or a proximal policy optimization (PPO) agent, which receives data from the sensor data acquisition module 10 and the physical state observation module 20, and outputs collaborative control actions to the collaborative control instruction generation and execution module 40.

[0036] In a specific embodiment using the DDPG algorithm, the reinforcement learning agent includes an actor network and a critic network. The actor network can be a fully connected neural network containing three hidden layers, whose input is a state space vector and whose output is a collaborative control action vector; the critic network's input is a state space vector and an action vector, and its output is the Q value of the state-action pair. The activation function of all hidden layers can adopt the rectified linear unit (ReLU) function.

[0037] In one embodiment, the reinforcement learning agent makes decisions based on an augmented state space, which is constructed as a vector containing the following three types of data components: The first category is real-time sensor data from the sensor data acquisition module 10, specifically including the current joint angle and angular velocity of the robot's carbon fiber joints, and the strain and temperature values ​​collected from the distributed fiber Bragg grating sensor array; The second category is the material physical state parameters calculated by the physical state observation module 20, specifically including the aforementioned creep damage measurement parameters and material microstructure stability parameters; The third category is the current material baseline physical parameter vector from the physical state observation module 20; By integrating these three types of data, the state space provides the reinforcement learning agent with comprehensive information about the external motion state and internal material health state.

[0038] The collaborative control action output by the reinforcement learning agent is an action vector containing two components. The first component is the torque adjustment action, which is defined as a scalar value and is used to generate an additional compensation torque; the second component is the physical stiffness adjustment action, which is defined as a scalar value and is used to set the target stiffness value of the controllable stiffness unit in the joint.

[0039] The training goal of the reinforcement learning agent is to maximize the long-term cumulative value of a compound reward function; Figure 4 In a specific embodiment, the compound reward function ( ) at each time step is calculated as follows: ; Where: For time; For in time Total reward value; For in time A task performance reward item, which is used to characterize the tracking accuracy of the joint motion trajectory; in one embodiment, the item is calculated as the negative square of the error between the expected joint trajectory and the actual joint trajectory; For in time A mid-term health maintenance reward term, which is used to penalize the degradation of the material physical state parameter; in one embodiment, the term is calculated as the negative value of the creep damage metric parameter value or its rate of change; For in time A long-term health gain reward term, which is used to penalize excessively rapid changes in the material baseline physical parameter vector; in one embodiment, this term is calculated as the negative value of the norm of the time change rate of the material baseline physical parameter vector; 、 ,and is a preset non-negative weight coefficient used to balance the relative importance of the three reward items, and the sum of the three is 1.

[0040] Refer to the attached Figure 1 and attached Figure 5 , step S4 is performed by Figure 1 The collaborative control instruction generation and execution module 40 shown is executed; this module receives the collaborative control action output by the reinforcement learning decision module 30, and parses it into two independent control instructions: an instruction torque and an instruction stiffness, and then drives the actuator of the robot's carbon fiber joint to execute synchronously.

[0041] Refer to the attached Figure 5 In a specific embodiment, the process of generating the command torque for driving the joint includes the following operations: First, the collaborative control instruction generation and execution module 40 calculates a predicted creep displacement based on the current material baseline physical parameter vector and stress and temperature states obtained from the physical state observation module 20 through a built-in short-term creep prediction model. This predicted creep displacement is superimposed on the original desired trajectory given by the upper-level motion planner to obtain a modified desired trajectory. This is a feedforward compensation operation. Next, the coordinated control instruction generation and execution module 40 provides the modified desired trajectory as input to a basic controller, which can be a proportional-integral-derivative (PID) controller or a linear quadratic regulator (LQR). The basic controller calculates a basic torque based on the deviation between the modified desired trajectory and the actual joint position.

[0042] Refer to the attached Figure 1 and attached Figure 5 , step S4 is performed by Figure 1 The collaborative control instruction generation and execution module 40 shown is executed; the collaborative control instruction generation and execution module 40 receives the collaborative control action from the reinforcement learning decision module 30, and generates physically executable instruction torque and instruction stiffness based on this, and finally sends these two instructions synchronously to the drive motor and controllable stiffness unit of the robot joint.

[0043] In a specific embodiment, the process of generating the command torque combines model-based feedforward compensation and reinforcement learning-based feedback regulation. Figure 5 , the process specifically includes: First, the collaborative control instruction generation and execution module 40 uses the creep strain ( ), the predicted creep displacement is calculated using the following error propagation model ( ): ; Where, is the predicted joint angle offset; is a coefficient related to the joint geometry; For in time Creep strain; For in time and temperature The relevant material elastic modulus.

[0044] In the calculation of the predicted creep displacement ( ), and then superimposed on the original expected trajectory given by the robot task planning layer ( ) to obtain a modified expected trajectory ( ), which is calculated as follows: ; The modified desired trajectory is fed into a basic controller, such as a proportional integral derivative (PID) controller, which calculates a basic torque based on the error between the modified desired trajectory and the trajectory fed back by the actual joint position sensor. This torque is used to complete the main trajectory tracking task. At the same time, the collaborative control instruction generation and execution module 40 converts the torque adjustment action component in the collaborative control action output by the reinforcement learning decision module 30 into a compensation torque with physical units ( ).

[0045] As a specific torque compensation method, the compensation torque ( ) can be calculated based on the predicted creep displacement ( ,Right now ), generated by a proportional derivative (PD) controller: ; Where, is the compensation torque; is the predicted creep displacement; is its rate of change; and are the preset proportional and differential gain coefficients.

[0046] The base torque output by the base controller is algebraically summed with the compensation torque to obtain the final command torque, which is then sent to the joint's drive motor for execution.

[0047] The process of generating the command stiffness is to map the physical stiffness adjustment action component in the collaborative control action output by the reinforcement learning decision module 30 to the actual physical stiffness value of the controllable stiffness unit. In a specific embodiment, the physical stiffness adjustment action is a value normalized to an interval; the collaborative control command generation and execution module 40 sets a minimum achievable stiffness value ( ) and the maximum stiffness value ( ).

[0048] The collaborative control instruction generation and execution module 40 uses the following linear mapping relationship to adjust the physical stiffness action value ( ) is converted to the final command stiffness ( ): ; Where: is the final generated instruction stiffness; is the minimum physical stiffness value that can be achieved by the controllable stiffness unit; is the maximum physical stiffness value achievable by the controllable stiffness unit; The action value is adjusted for the physical stiffness within the value range received from the reinforcement learning decision module 30.

[0049] In a specific embodiment, when the controllable stiffness unit is a magnetorheological damper, the coordinated control instruction generation and execution module 40 further includes a current driving unit, which calculates the command stiffness ( ) is converted into a corresponding driving current value according to the pre-calibrated stiffness and current characteristic curve, and applied to the electromagnetic coil of the magnetorheological damper.

[0050] This calculated command stiffness value is sent to the controller of the controllable stiffness unit to adjust its physical stiffness.

[0051] Refer to the attached Figure 1 , step S5 is performed by Figure 1 The baseline parameter update module 50 shown is executed; the function of the baseline parameter update module 50 is to perform online, closed-loop updates on the material baseline physical parameter vector stored inside the physical state observation module 20 based on the actual load history of the robot's carbon fiber joints during long-term service, so as to reflect the evolution of the material's physical properties due to factors such as aging and fatigue.

[0052] The baseline parameter update module 50 is connected to the collaborative control instruction generation and execution module 40 to receive the execution history data of the instruction torque and instruction stiffness. The module contains a data storage unit for continuously recording the stress history and temperature history of the joint, wherein the stress history is calculated by combining the received instruction torque history with the known joint geometric model; the temperature history is directly derived from the sensor data acquisition module 10.

[0053] In a specific embodiment, the parameter update is performed periodically; the baseline parameter update module 50 sets a preset time window, for example, every 100 working hours. During this time window, the module continues to accumulate the stress and temperature data of the joint, and when a time window ends, an update calculation is triggered.

[0054] The core of the update calculation is based on an evolution model of the material's baseline physical parameter vector, which describes the relationship between the rate of change of the parameter vector and its load history and current state. In one embodiment, the evolution model is expressed as: ; Where: For time; For in time The material baseline physical parameter vector; For in time The rate of change of the material baseline physical parameter vector; It is a preset evolution function, which is a vector function. Its specific form is constructed based on prior knowledge of materials science. For example, for the temperature-related components in the parameter vector, its evolution can be described by the Arrhenius model to describe the effect of temperature-accelerated aging; for the stress-related components, its evolution can be described by a model based on damage mechanics or fatigue accumulation. Deadline The joint stress history is the stress time series recorded by the baseline parameter updating module 50; Deadline The joint temperature history is the temperature time series recorded by the baseline parameter updating module 50 .

[0055] For example, for the material baseline physical parameter vector Creep coefficient in Component, its evolution function A specific form of can be expressed as: ; Where, is the rate of change of creep coefficient; is the material constant calibrated by experiment; and are the average stress and average absolute temperature in the last update cycle respectively; is the aging activation energy of the material; is the ideal gas constant. Through this model, the creep coefficient evolution increment over time can be quantitatively calculated based on the recorded stress and temperature history.

[0056] Each time an update is triggered, the baseline parameter update module 50 takes the recorded stress history and temperature history within the time window, as well as the current material baseline physical parameter vector as input, substitutes them into the above-mentioned evolution model formula, and calculates the increment of the parameter vector within the time window through a numerical integration method (for example, the first-order Euler method or the fourth-order Runge-Kutta method). Then, the increment is added to the old parameter vector to obtain the updated new parameter vector.

[0057] After the calculation is completed, the baseline parameter updating module 50 writes the updated material baseline physical parameter vector into the designated storage area of ​​the physical state observation module 20 for use in the state observation calculation in the next working cycle.

[0058] Refer to the attached Figure 1 -Attached Figure 5 To more fully illustrate the collaborative working process of the technical solution of the present invention, we will use a specific application scenario as an example. This scenario involves a six-axis robot equipped with the control system of the present invention, with its carbon fiber wrist joint performing a cyclic, high-load handling task: continuously moving a heavy object back and forth between points A and B.

[0059] At the initial stage of task execution, the robot joints are in a brand new state. When the robot starts its first handling cycle, the system works as follows: When the joint begins to bear the load and move, the sensor data acquisition module 10 collects the strain and temperature data of each measuring point inside the joint in real time. In the initial state, the temperature is the ambient temperature, and the total strain is mainly elastic strain; The physical state observation module 20 receives this data, and its internal stress-strain decoupling unit calculates the stress distribution inside the joint. The creep damage metric parameter calculated by the mid-term health state calculation unit is initially set to zero, and the material baseline physical parameter vector in the long-term health state calculation unit is the initial calibration value. The reinforcement learning decision module 30 takes as input a state vector containing the current joint angle, angular velocity, zero-damage metric, and initial baseline parameters. Since the material health state parameters are all at optimal values, the penalty values ​​for the medium-term health maintenance reward and the long-term health gain reward in the composite reward function are both zero, and the task performance reward now dominates. Therefore, the coordinated control actions output by the reinforcement learning agent will primarily aim to maximize trajectory tracking accuracy. For example, a torque adjustment action for generating a large acceleration torque and a physical stiffness adjustment action set to a medium value will be output. Based on this action, the collaborative control instruction generation and execution module 40 generates an instruction torque sufficient to drive the joint to move quickly and an instruction stiffness of medium size, and executes them synchronously.

[0060] After the robot works continuously for several hours, the temperature of its joints rises due to heat generated by the motor and the movement of the molecular chains within the material. At the same time, continuous operation under high stress causes the material to creep. At this point, the system's workflow demonstrates adaptive adjustment: The sensor data acquisition module 10 detects a significant increase in the internal temperature of the joint and can still detect a small residual strain when the joint is unloaded (for example, when a heavy object is dropped at point A or point B). After the physical state observation module 20 receives the new sensor data, its mid-term health state calculation unit 22 calculates a non-zero and continuously increasing creep damage metric parameter based on the integral of the stress history and creep strain rate; When the reinforcement learning decision module 30 receives a state vector containing an increasing temperature and a non-zero creep damage metric parameter, the medium-term health maintenance reward term in the composite reward function begins to generate a negative reward value. To maximize the long-term cumulative reward, the agent's strategy is adjusted to output a new coordinated control action. For example, this action may correspond to a slight negative compensation torque to smooth the motion curve and reduce peak stress, and a larger physical stiffness adjustment action value to enhance the physical support of the joint and resist creep deformation. The collaborative control instruction generation and execution module 40 generates an instruction torque with a slightly lower peak value according to this new action, and at the same time generates a higher instruction stiffness, which is sent to controllable stiffness units such as magnetorheological dampers for execution; in this way, the system actively slows down the damage accumulation rate of the material at the expense of minimal task execution speed.

[0061] When the system has been running for a preset long time (e.g., 1000 hours), the baseline parameter update module 50 is triggered to perform an online update; The module reads the complete stress and temperature history data stored internally over the past 1000 hours. It then substitutes this historical data, along with the current material baseline physical parameter vector, into a preset evolution model formula for the material baseline physical parameter vector. The calculated results reflect changes in the material's physical properties after long-term service, such as a slight increase in the creep coefficient or a slight decrease in the elastic modulus. The baseline parameter updating module 50 writes the calculated new material baseline physical parameter vector into the storage area of ​​the physical state observation module 20 .

[0062] After this update is completed, the entire system continues to perform its tasks, but its internal model has adapted to the aging of the material. In subsequent calculations, the physical state observation module 20 will use this updated parameter vector, which is closer to the actual physical condition of the material, to more accurately decouple strain and predict creep. This makes the decision basis of the reinforcement learning decision module 30 more accurate and the feedforward compensation of the collaborative control instruction generation and execution module 40 more precise, thus ensuring that the robot can maintain optimal task performance and joint health balance at different stages throughout its life cycle.

[0063] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A dynamic load collaborative control method for robot carbon fiber joints based on reinforcement learning, characterized in that: The method comprises the following steps: S1. Acquire sensor data in the carbon fiber joint of the robot, wherein the sensor data includes strain and temperature; S2. Calculating, based on the sensor data, material physical state parameters of the joint material in a mid-term health state and a material baseline physical parameter vector of the joint material in a long-term health state using a physical state observation model; S3. Using the sensor data, the material physical state parameters, and the material baseline physical parameter vector as state inputs of a reinforcement learning agent, and having the reinforcement learning agent output a coordinated control action including a torque adjustment action and a physical stiffness adjustment action; S4. generating a command torque for driving the joint and a command stiffness for adjusting a controllable stiffness unit in the joint according to the coordinated control action, and synchronously executing the command torque and command stiffness on the joint; S5. Update the material baseline physical parameter vector in the physical state observation model according to the execution history of the command torque and the command stiffness, and provide the updated material baseline physical parameter vector to the physical state observation model for subsequent calculations.

2. The method for collaborative control of dynamic loads of carbon fiber joints of robots based on reinforcement learning according to claim 1 is characterized in that: In step S1, the step of obtaining sensor data in the carbon fiber joint of the robot includes: The strain and temperature are collected by a distributed fiber grating sensor array embedded in the carbon fiber joint structure of the robot.

3. The method for collaborative control of dynamic loads of carbon fiber joints of robots based on reinforcement learning according to claim 1, characterized in that: In step S2, the material physical state parameters of the joint material in the mid-term health state are calculated by the physical state observation model, including: Creep damage metric, used to quantify the cumulative energy dissipation of the material due to creep; The material microstructure stability parameter is used to characterize the stability of the material's crystallinity and amorphous regions.

4. The method for collaborative control of dynamic loads of carbon fiber joints of robots based on reinforcement learning according to claim 3 is characterized in that: The calculation steps of the creep damage metric parameters include: Calculating the stress and creep strain rate inside the joint based on the sensing data; The creep damage metric parameter is calculated based on the stress and creep strain rate. The real-time change rate of the creep damage metric parameter satisfies the following formula: ; Where, For time; is the real-time change rate of creep damage measurement parameters; is the stress inside the joint; is the creep strain rate.

5. The method for collaborative control of dynamic loads of carbon fiber joints of robots based on reinforcement learning according to claim 1, characterized in that: In step S2, the material baseline physical parameter vector of the long-term health state of the joint material at least includes a creep coefficient, a time index, and an initial elastic modulus for describing the creep compliance of the material.

6. The method for collaborative control of dynamic loads of carbon fiber joints of robots based on reinforcement learning according to claim 1, characterized in that: Step S3 specifically includes: The reinforcement learning agent is trained by maximizing the long-term cumulative value of a compound reward function, which includes: Task performance reward item used to characterize the accuracy of joint motion trajectory tracking; A mid-term health maintenance bonus item, wherein the mid-term health maintenance bonus item is negatively correlated with the degree of deterioration of the material's physical state parameters; A long-term health gain reward item, wherein the long-term health gain reward item is negatively correlated with the time rate of change of the material baseline physical parameter vector.

7. The method for collaborative control of dynamic loads of carbon fiber joints of robots based on reinforcement learning according to claim 1, characterized in that: The controllable stiffness unit in the joint in step S4 is a magnetorheological damper, an electrorheological damper or a piezoelectric stack actuator.

8. The method for collaborative control of dynamic loads of carbon fiber joints of robots based on reinforcement learning according to claim 1, characterized in that: In step S4, the step of generating a command torque for driving the joint includes: The predicted creep displacement is superimposed on the expected trajectory to obtain a corrected expected trajectory; generating a base torque according to the modified desired trajectory by a proportional-integral-derivative base controller; The basic torque is added to the compensation torque generated according to the torque adjustment action to obtain the command torque.

9. The method for collaborative control of dynamic loads of carbon fiber joints of robots based on reinforcement learning according to claim 1, characterized in that: In step S5, the step of updating the material baseline physical parameter vector in the physical state observation model includes: The update of the material baseline physical parameter vector is based on the following evolution model: ; Where, For time; is the rate of change of the material baseline physical parameter vector; is the preset evolution function; is the stress history of the joint; Temperature history for joints; For in time The material baseline physical parameter vector; The evolution function is a function based on the Arrhenius model, which is used to characterize the accelerated effect of temperature on material aging.

10. A robot carbon fiber joint dynamic load collaborative control system based on reinforcement learning, applied to the method according to any one of claims 1 to 9, characterized in that: The system comprises: A sensor data acquisition module is used to obtain sensor data in the carbon fiber joints of the robot, wherein the sensor data includes strain and temperature; a physical state observation module, configured to receive the sensor data and calculate, based on the sensor data, material physical state parameters of the joint material in the mid-term health state and a material baseline physical parameter vector of the joint material in the long-term health state; a reinforcement learning decision module, configured to receive the sensor data, the material physical state parameters, and the material baseline physical parameter vector as state inputs, and output a coordinated control action including a torque adjustment action and a physical stiffness adjustment action; a collaborative control instruction generation and execution module, configured to generate and synchronously execute an instruction torque for driving the joint and an instruction stiffness for adjusting a controllable stiffness unit in the joint according to the collaborative control action; The baseline parameter updating module is used to update the material baseline physical parameter vector in the physical state observation module according to the execution history of the command torque and the command stiffness, and provide the updated material baseline physical parameter vector to the physical state observation module.

Citation Information

Patent Citations

  • Robot state prediction method and device, electronic equipment and storage medium

    CN117733874A

  • Industrial robot control system and method

    CN120255424A

  • Pneumatic soft-bodied robot capable of autonomously advancing through orifice and control method of pneumatic soft-bodied robot

    CN120514303A

  • Robot joint

    CN218947725U

  • Systems, device and object comprising electroactive polymer material, methods and uses relating to operation and provision thereof

    WO2009038501A1

Cited By

  • Flexible robot deformation control and environment adaptation system based on deep learning

    CN122077641A