Robot whole-body motion control method and system based on reinforcement learning
Through a robot full-body motion control method based on reinforcement learning, using a deep reinforcement learning model and Ethernet control automation technology bus, combined with a three-layer protection mechanism, the generalization and real-time problems of traditional robot control systems in complex scenarios are solved, high-frequency real-time control and multi-level security protection are achieved, and the reliability and stability of the robot system are improved.
Patent Information
- Application Number
- CN202510749248.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional robot control systems have poor generalization and slow real-time response when faced with complex and changing scenarios, and their safety protection mechanisms are single and cannot effectively deal with various abnormal situations such as sensor noise and communication failures, resulting in poor reliability and stability of the robot in actual operation.
A reinforcement learning-based robot whole-body motion control method is adopted. Inertial measurement unit data and motor feedback data are collected in real time through inference nodes. A deep reinforcement learning model accelerated by TensorRT is used to generate control instructions. The Ethernet control automation technology bus is combined to issue instructions, and a three-layer protection mechanism is constructed at the electrical layer, mechanical layer, and software layer.
It realizes high-frequency real-time control of the robot's whole-body movement, multi-scenario adaptive adjustment and multi-level safety protection, improves the generalization and safety of the robot system, reduces the overall response delay, and enhances the reliability and stability of the robot system.
Smart Images

Figure CN120735003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot control systems, and in particular to a robot whole-body motion control method and system based on reinforcement learning. Background Art
[0002] With the rapid development of artificial intelligence, sensing technology, and automatic control technology, robotic systems are being widely used in a variety of fields, including industrial manufacturing, medical assistance, service industries, military reconnaissance, and domestic services. As the core bridge between the robot body and its intelligent behavior, the robotic control system directly affects its motion accuracy, response speed, environmental adaptability, and task execution efficiency. Traditional robotic control systems are mostly based on classical control theories such as PID control and fuzzy control, and are suitable for controlling fixed trajectories in structured environments. However, with the increasing complexity of application scenarios, modern robotic systems have placed higher demands on the real-time, adaptable, and intelligent control systems.
[0003] Traditional robot control solutions have the following shortcomings. First, the control algorithm has poor generalization and relies on complex rules set in advance by humans. It is difficult to automatically adjust joint control instructions according to the real-time environment, and performs poorly when faced with complex and changing scenarios (such as different terrains and dynamic loads). Second, the real-time response speed is slow and cannot meet high-frequency control requirements. It is difficult to cope with situations where the robot moves quickly or needs to adjust its movements in a timely manner. In addition, the safety protection mechanism is single, and most of the time it only provides simple protection from one aspect of hardware or software. It lacks a multi-level and multi-dimensional collaborative protection system and cannot effectively deal with various abnormal situations such as sensor noise and communication failures. As a result, the robot's reliability and stability are poor in actual operation. Summary of the Invention
[0004] The main purpose of the present invention is to provide a robot whole-body motion control method and system based on reinforcement learning, aiming to solve the technical problems raised in the above background technology.
[0005] The present invention proposes a robot whole-body motion control method based on reinforcement learning, comprising the following steps: Real-time collection of inertial measurement unit data and motor feedback data through inference nodes; Inputting the inertial measurement unit data and the motor feedback data into a deep reinforcement learning model accelerated by TensorRT to generate motor control instructions; The control instructions are issued by the master node using the Ethernet control automation technology bus, wherein the upper layer of the master node runs a four-state state machine including an initialization state, an enable state, a motor control state, and a fault state; the lower layer implements, through a protocol stack, the following: converting the control instructions into service data object data packets, using a distributed clock synchronization mechanism to achieve clock synchronization, and real-time acquisition and updating of slave device status words; The motor drive node parses the service data object data packet and drives the motor to realize the robot joint motion control; A three-layer protection mechanism is implemented during the above steps, including electrical layer protection, mechanical layer protection and software layer protection.
[0006] Preferably, the initialization state of the state machine includes: Perform hardware self-test, check emergency stop signal and Ethernet control automation technology bus ring network connection status, and verify the number of slave stations and configuration matching; Initialize parameters, load software configuration parameters and control parameters, including joint motion range threshold and current protection threshold; Establishing a communication link, completing the handshake protocol between the Ethernet control automation technology bus master station and the slave station, and determining that the connection is successful when the packet loss rate is less than a preset threshold; Initialize the state transition condition, automatically jump to the enabled state after the self-test passes, and enter the fault state if the self-test fails.
[0007] Preferably, the enabled state of the state machine includes: The soft start process uses a progressive torque loading strategy to increase the initial torque from low to high according to the preset time interval until full load is reached; Communication quality verification: real-time monitoring of communication packet loss rate. If the packet loss rate exceeds the threshold alarm value within a preset period, a fault state is triggered; Activate the watchdog timer and set a detection period. If no heartbeat signal is received within the detection period, perform an emergency shutdown. The state transition condition is enabled, and the state automatically jumps to the motor control state after the soft start is completed and the communication quality is stable.
[0008] Preferably, the steps of implementing the three-layer protection mechanism include: Dynamically adjust the limit range according to the real-time movement speed of the joint; When it is detected that the joint position exceeds the dynamic limit range, a joint over-limit exception is triggered and damping control is performed; Set the proportional gain to zero and apply a high derivative gain.
[0009] Preferably, the steps of implementing the three-layer protection mechanism include: Adopt sliding window detection method to store multiple frames of continuous motor angle data; When the angle difference between two adjacent frames exceeds the preset arc, the second-order verification is started to check whether the direction of the torque change rate matches the motion trend; If the data is abnormal within the preset frame period, an emergency stop is triggered; Set basic timeout thresholds and calculate network jitter coefficients; Dynamically generate a timeout threshold based on the basic timeout threshold and the calculated network jitter coefficient; When the instruction response time exceeds the actual timeout threshold, a protection action is triggered; Real-time monitoring of the current motor temperature and ambient temperature; When the temperature change rate is greater than the heat dissipation coefficient multiplied by the difference between the current temperature and the ambient temperature, an overheating warning is triggered.
[0010] The present invention also discloses a robot whole-body motion control system based on reinforcement learning, comprising: An inference node, configured to collect inertial measurement unit data and motor feedback data, and to input the inertial measurement unit data and the motor feedback data into a deep reinforcement learning model accelerated by TensorRT to generate motor control instructions; The master node is configured to issue the control instructions via the Ethernet control automation technology bus, wherein the upper layer of the master node runs a four-state state machine including an initialization state, an enable state, a motor control state, and a fault state; and the lower layer implements, through a protocol stack, the following: converting the control instructions into service data object data packets, implementing clock synchronization using a distributed clock synchronization mechanism, and collecting and updating slave device status words in real time; A motor driving node is used to parse the service data object data packet and drive the motor to realize the robot joint motion control; The protection module is used to implement a three-layer protection mechanism, including electrical layer protection, mechanical layer protection and software layer protection, which are respectively realized through the electrical layer protection unit, mechanical layer protection unit and software layer protection unit. The electrical layer protection unit is used to integrate the dynamic heartbeat detection circuit and the overcurrent protection relay; the mechanical layer protection unit is used to install the hard limit switch and the buffer zone detection sensor; the software layer protection unit is used to deploy the dynamic soft limit algorithm and the multi-dimensional anomaly detection program.
[0011] Preferably, the master station node includes: Initialization status unit, performs hardware self-test and parameter initialization, and activates the fault state when a mismatch in the number of slaves is detected; Enable state unit, used to output torque according to the gradient increasing strategy; and used to calculate the packet loss rate in real time and control the watchdog timer; The motor control state unit is used to support the expansion of motion sub-states, including independent control logic for walking, standing, and jumping; and is used to plan joint motion curves based on deep reinforcement learning output instructions; The fault status unit is used to distinguish between hardware faults and software faults. Hardware faults include overcurrent and overtemperature, and software faults include communication timeout and data anomalies; and is used to execute joint position locking or system safety shutdown.
[0012] Preferably, the software layer protection unit includes: A dynamic limit calculator, used to receive joint velocity data in real time, calculate velocity scaling factors, and output dynamic limit thresholds; Damping controller, used to switch control mode when over-limit is triggered; The anomaly detector is used to store the motor angle data, compare the logical relationship between the angle change and the torque change rate, and generate an overheating warning signal based on the heat dissipation equation.
[0013] The present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, it implements the steps of a robot whole-body motion control method based on reinforcement learning.
[0014] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the computer program implements the steps of a robot whole-body motion control method based on reinforcement learning.
[0015] The beneficial effects of the present invention are as follows: the present invention constructs a hierarchical control process of "inference node collecting and analyzing data - deep reinforcement learning model generating instructions - master station node scheduling and issuing instructions - motor drive node execution control - three-layer protection mechanism to ensure safety", uses a deep reinforcement learning model accelerated by TensorRT to replace the traditional control algorithm, and combines the hard real-time communication capabilities of the Ethernet control automation technology bus to achieve high-frequency real-time control of the robot's whole-body movement, multi-scenario adaptive adjustment and multi-level safety protection, and solves the problems of poor generalization, slow real-time response speed and single safety protection mechanism in traditional robot control. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a schematic diagram of a method flow chart according to an embodiment of the present application.
[0017] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0018] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0019] like Figure 1 As shown, the present application provides a robot whole-body motion control method based on reinforcement learning, comprising: S1, real-time collection of inertial measurement unit data and motor feedback data through inference nodes. The inertial measurement unit data includes angular velocity and gravity projection, and the motor feedback data includes motor position and motor speed. The data collection cycle does not exceed 1 millisecond. The inertial measurement unit data is parsed and forwarded through real-time dynamic positioning drive nodes. S2, inputting the inertial measurement unit data and the motor feedback data into a deep reinforcement learning model accelerated by TensorRT to generate motor control instructions; S3, issuing the control instructions via the master node using the Ethernet control automation technology bus, wherein the upper layer of the master node runs a four-state state machine including an initialization state, an enable state, a motor control state, and a fault state; the lower layer, through the protocol stack, implements: converting the control instructions into service data object data packets, using a distributed clock synchronization mechanism to achieve clock synchronization with a period of no more than 1 millisecond, and real-time acquisition and updating of slave device status words; S4, parsing the service data object data packet and driving the motor through the motor drive node to realize robot joint motion control; S5. During the execution of the above steps S1-S5, a three-layer protection mechanism is implemented, including electrical layer protection, mechanical layer protection and software layer protection.
[0020] Electrical layer protection: Dynamic heartbeat detection is performed, with a baseline detection period of 500 milliseconds and a tolerance window of ±20%. When the current is detected to exceed 150% of the rated value, the metal oxide semiconductor field effect transistor is triggered to shut down; Mechanical layer protection: Set hard limits to constrain the range of joint motion, and set a warning buffer 2 degrees before the hard limit trigger point; Software layer protection: Implement dynamic soft limit control, damping control and multi-dimensional anomaly detection.
[0021] As described in the above steps S1-S5, the present invention constructs a hierarchical control process of "inference node collects and analyzes data - deep reinforcement learning model generates instructions - master station node schedules and issues instructions - motor drive node executes control - three-layer protection mechanism ensures safety", uses the deep reinforcement learning model accelerated by TensorRT to replace the traditional control algorithm, and combines the hard real-time communication capabilities of the Ethernet control automation technology bus to achieve high-frequency real-time control of the robot's whole-body movement, multi-scenario adaptive adjustment and multi-level safety protection, and solves the problems of poor generalization of complex scenarios, slow real-time response speed and single safety protection mechanism in traditional robot control.
[0022] During robot motion, the robot must integrate the angular velocity and gravity projection data provided by the inertial measurement unit (IMU) with the position and velocity data fed back by the motors in real time to precisely control the movements of each joint. Traditional PID-based control algorithms rely on manual parameter tuning, making them difficult to quickly adapt to complex scenarios such as terrain changes and sudden load fluctuations (for example, the robot must dynamically adjust its gait when navigating stairs, slopes, and other different surfaces). Furthermore, mechanical systems require extremely high real-time performance. Delays exceeding 10 milliseconds in issuing control commands can lead to unstable motion or even loss of control (for example, a bipedal robot could fall due to a loss of balance due to delays). At the hardware level, motor overcurrent (such as a sudden surge in current when a joint becomes stuck) can burn out the driver circuitry, and joint movement exceeding physical limits can damage the mechanical structure. At the software level, abnormal sensor data or network communication delays can lead to miscontrol. Therefore, a multi-layered protection system is necessary at the electrical, mechanical, and software levels to filter interference and respond promptly to anomalies to prevent equipment damage or escalate failures.
[0023] To address the aforementioned issues, traditional control algorithms suffer from poor generalization, insufficient real-time performance, and limited and passive safety protection. However, the present invention replaces the traditional PID algorithm with a deep reinforcement learning model. Through offline model training, the model learns a variety of motion patterns (such as walking, standing, and jumping), enabling the robot to automatically adjust joint control commands based on the real-time environment, eliminating the need for manual pre-setting of complex rules. Furthermore, the present invention utilizes the Ethernet control automation technology bus to rapidly issue commands, and combines it with TensorRT to accelerate deep reinforcement learning model inference, resulting in lower overall response latency, a significant improvement over traditional solutions, ensuring high-frequency, real-time control of the robot's joints. Furthermore, a triple protection mechanism is constructed at the electrical, mechanical, and software levels: the electrical layer implements overcurrent protection and heartbeat detection through hardware circuits; the mechanical layer prevents joint overtravel collisions through physical limits and buffer warning zones; and the software layer implements active protection and rapid response through algorithms that dynamically adjust limit ranges, apply damping control, and perform multi-dimensional anomaly detection.
[0024] Specifically, the inference node collects inertial measurement unit (IMU) data (angular velocity, gravity projection) and motor feedback data (position, velocity) in real time through a hardware interface, with a data collection cycle of less than 1 millisecond. After collection, the IMU data is parsed by the real-time dynamic positioning drive node and forwarded to subsequent processing stages, ensuring real-time and accurate data transmission. This high-frequency acquisition captures microsecond-level motion changes in robot joints (such as the sudden change in ankle angular velocity during a jump), providing real-time and accurate state input for deep reinforcement learning models. For example, in a robot with 30 joints, data can be collected 1000 times per second, each containing 210 dimensions of information (5-dimensional inertial data + 14-dimensional motor feedback data x 30 joints), ensuring that the model can perceive the robot's entire body state in a timely manner. The hardware-accelerated parsing capabilities of the real-time dynamic positioning drive node avoid data processing delays, ensuring that front-end data can be quickly incorporated into subsequent decision-making processes.
[0025] By inputting the collected inertial measurement unit data and motor feedback data into a deep reinforcement learning model accelerated by TensorRT, the model generates motor control commands based on the motion control strategy learned through pre-training. TensorRT technology can optimize and accelerate the model, ensuring that the command generation cycle does not exceed 2 milliseconds. The deep reinforcement learning model can autonomously learn complex joint motion coordination strategies through extensive simulation training (such as simulating various terrains in a virtual environment). For example, in an obstacle crossing scenario, the model can directly output coordinated motion commands for the entire body's joints, avoiding the delays associated with step-by-step gait planning and real-time adjustment in traditional methods. TensorRT's acceleration capabilities enable the model to achieve fast inference on embedded hardware (such as the Jetson series), meeting the computing power requirements of real-time robot control.
[0026] The upper layer of the master node runs a four-state state machine consisting of initialization, enable, motor control, and fault states. The initialization state performs hardware self-tests (such as checking Ethernet control automation technology bus ring network connectivity and verifying the number of slaves and configuration compatibility) and initializes parameters (such as setting joint motion range thresholds and current protection thresholds). The enable state starts the motors through progressive torque loading and activates a watchdog timer to monitor communication status. The motor control state controls the robot's movements, such as walking and standing, based on instructions generated by a deep reinforcement learning model. The fault state initiates emergency shutdowns and other actions when hardware faults (such as overcurrent) or software faults (such as communication timeouts) are detected. The lower layer of the master node converts control commands into service data object packets through the protocol stack. A distributed clock synchronization mechanism (period ≤ 1 millisecond) ensures clock consistency among all slave devices and collects slave status words in real time to monitor device operating status. This state machine mechanism ensures the orderliness and reliability of the robot's control process. For example, the initialization state undergoes rigorous hardware self-tests and parameter verification to prevent subsequent control failures due to device connection anomalies or parameter errors. Progressive torque loading in the enable state prevents mechanical shock during motor startup (for example, the initial torque starts at 10% of the rated value and increases by 5% every 50 milliseconds to full load), extending device life. Distributed clock synchronization within the Ethernet control automation technology bus ensures synchronized movement of multiple motor drive nodes, enabling the left and right leg joints to swing precisely and in coordination along a predetermined trajectory during walking, avoiding gait disturbances caused by clock deviations.
[0027] The motor drive node receives service data object data packets, parses them, and generates pulse-width modulated signals to drive the motors, achieving motion control of the robot joints. During the drive process, the motor status is collected in real time and fed back to the inference node, forming a closed control loop. High-frequency command execution ensures the accuracy of the robot's joint movements. For example, in a precise grasping scenario involving a robotic arm, the motor can quickly adjust its speed and position based on real-time commands, achieving millimeter-level positioning accuracy. Feedback data such as current and temperature provides the basis for software-level anomaly detection. For example, if an abnormal increase in motor current is detected, the software-level protection mechanism can promptly trigger a protective action.
[0028] The electrical layer of the three-layer protection mechanism uses dynamic heartbeat detection. For example, the master node sends a heartbeat signal to the slave node at a base period of 500 milliseconds (allowing a ±20% tolerance). If no response is received within the timeout, the metal oxide semiconductor field effect transistor is immediately triggered to shut down, cutting off the motor power supply. This prevents motor loss of control due to communication interruption (such as a drone motor continuing to rotate at high speed after losing the control signal). If the motor current exceeds 150% of the rated value (such as a joint stuck by foreign objects), the drive circuit is shut down within microseconds to prevent overheating and damage to the motor or driver board.
[0029] Mechanical protection includes hard limits, which are physical limiters (such as photoelectric switches) installed at the extremes of joint motion. When a joint approaches its limit, further movement is mechanically blocked to prevent excessive rotation or swinging due to software control failure (for example, damage to a shoulder joint due to incorrect control rotation exceeding the physically permitted range). It also includes a warning buffer zone, set two degrees before the hard limit trigger point. When the joint enters this zone, the software layer triggers damping control in advance, slowing the joint's motion and reducing the impact of hitting the hard limit (for example, a robot arm automatically decelerates when approaching the edge of a worktable to avoid a violent collision).
[0030] Software-level protection includes dynamic soft limit control, which dynamically adjusts the limit range based on the real-time joint movement speed. The faster the speed, the smaller the limit range. For example, when the joint speed reaches the maximum allowable speed, the limit range is automatically reduced by 20% to prevent inertial overtravel during high-speed movement (for example, the knee joint swing amplitude automatically narrows when the robot is running, reducing the risk of falls). It also includes damping control. When the joint position exceeds the dynamic soft limit range, the proportional gain is immediately set to zero and a high differential gain is applied to quickly reduce the joint movement speed to avoid mechanical impact. This control strategy can shorten the joint braking time from the traditional 200 milliseconds to less than 50 milliseconds, effectively reducing collision damage. It also includes multi-dimensional anomaly detection, using a sliding window to detect the difference in motor angles between two adjacent frames. For example, if it exceeds 0.6 radians, it further verifies whether the torque change rate matches the motion trend to avoid false shutdowns caused by sensor noise (such as occasional angle jumps but the torque does not change synchronously, which is judged as invalid data); dynamically adjusts the timeout threshold according to network jitter. For example, when the network delay increases, the tolerance threshold of the command response time is automatically extended to avoid unnecessary protection actions caused by temporary network fluctuations; and monitors the motor temperature change rate in real time. When the temperature change is abnormal (greater than the heat dissipation coefficient multiplied by the difference between the current temperature and the ambient temperature), it triggers an overheating warning in advance, prompting the system to reduce load or shut down for heat dissipation.
[0031] In summary, the present invention adopts a deep reinforcement learning model, combines the Ethernet control automation technology bus and TensorRT acceleration technology, and constructs a triple protection mechanism. Compared with traditional solutions, it can improve the generalization of the robot control algorithm, realize high-frequency real-time control, and significantly reduce the overall response delay; at the same time, the constructed multi-level active safety protection system can quickly respond to abnormal situations, enhancing the safety and reliability of the robot system.
[0032] In one embodiment of the present invention, the initialization state of the state machine includes: Perform hardware self-test, check emergency stop signal and Ethernet control automation technology bus ring network connection status, and verify the number of slave stations and configuration matching; Initialize parameters, load software configuration parameters and control parameters, including joint motion range threshold and current protection threshold; Establishing a communication link, completing the handshake protocol between the Ethernet control automation technology bus master station and the slave station, and determining that the connection is successful when the packet loss rate is less than a preset threshold; Initialize the state transition condition, automatically jump to the enabled state after the self-test passes, and enter the fault state if the self-test fails.
[0033] As described in the above steps, the present invention achieves the goal of ensuring that the robot system hardware is normal, the parameters are correctly configured, and the communication connection is stable by performing a series of orderly operations in the initialization state, including hardware self-test, parameter initialization, and communication link establishment, thereby laying the foundation for the subsequent robot to enter the enabled state and operate normally.
[0034] After starting or resetting the robot, it is essential to confirm that the system hardware, parameter settings, and communication links are in normal condition for safe and reliable operation. Hardware failures (such as abnormal emergency stop signals or interrupted bus connections) can lead to uncontrollable and dangerous robot operations. Parameter misconfigurations (such as improperly set joint range thresholds or current protection thresholds) can cause abnormal robot movement or hardware damage. An unstable communication link (such as excessive packet loss) can prevent the accurate transmission and execution of control commands. Therefore, the goal of the initialization state is to proactively identify and resolve these potential issues.
[0035] The traditional robot initialization process may have incomplete hardware self-checks, only checking some key hardware and missing some potential fault points; parameter configurations are mostly fixed values, lacking flexibility and difficult to adapt to the needs of different application scenarios; when the communication link is established, the verification method for connection stability is simple and cannot effectively deal with problems such as packet loss and delay in complex network environments, which can easily lead to unstable connections or misjudgment of connection success.
[0036] Through systematic and refined process design, the present invention conducts comprehensive self-inspection of hardware, covering key factors such as emergency stop signals, bus connection status, and the number of slave stations; during parameter initialization, software configuration parameters and control parameters can be flexibly loaded; in terms of establishing communication links, a strict handshake protocol is adopted, and the packet loss rate is used as the quantitative judgment standard for connection success, ensuring the reliability of the communication link and comprehensively solving the problems existing in traditional solutions.
[0037] Specifically, through appropriate hardware interfaces and detection circuits, the system checks the emergency stop signal (determining whether the emergency stop button has been pressed and signal transmission is normal), monitors the Ethernet control automation technology bus ring network connection (checking for loose network cables and whether nodes are communicating properly), and verifies the compatibility of the number of slaves with the configuration (comparing the actual number of connected slave devices with the pre-configured number). This data is collected in real time via hardware sensors and communication links. Comprehensive hardware self-checks can detect potential hardware faults early in the robot's startup, preventing hardware issues from causing robot loss of control or damage during operation. For example, an abnormal emergency stop signal may prevent the robot from stopping in time, potentially causing a safety incident. Poor bus connectivity or an incompatible number of slaves can affect the transmission and execution of control commands, leading to abnormal robot movement. Preemptive detection and resolution of these issues ensures safe and stable robot operation.
[0038] By loading pre-set software configuration parameters and control parameters from a storage device (such as EEPROM, flash memory, etc.), including the joint motion range threshold (limiting the angular range of the joint movement) and the current protection threshold (setting the maximum current value allowed for normal motor operation), etc., correct parameter initialization is an important guarantee for the normal operation of the robot. Reasonable setting of the joint motion range threshold can prevent the joint from moving beyond the limit position and causing mechanical damage; accurate setting of the current protection threshold can ensure that protective measures can be taken in time when the motor has an overcurrent condition to prevent the motor from burning out. For example, if the joint motion range threshold is set too large, the robot joint may collide with surrounding objects; if the current protection threshold is set too low, the motor may frequently shut down due to overload protection due to misjudgment, affecting the normal operation of the robot; if it is set too high, the motor cannot be effectively protected.
[0039] Communication between the master and slave nodes is established through a handshake protocol based on the Ethernet control automation technology bus. During communication, the packet loss rate of data transmission is monitored in real time. When the packet loss rate is less than a preset threshold (such as 0.1%, which can be adjusted according to the actual application scenario), the connection is determined to be successful. A reliable communication link is the foundation for precise robot control. A strict handshake protocol and a connection determination mechanism based on packet loss rate ensure stable and accurate data transmission between the master and slave nodes. For example, during robot motion, if the packet loss rate of the communication link is too high, control instructions may not be accurately transmitted to the slave motor drive node, resulting in deviation or loss of control of the robot's motion. This step can effectively avoid such problems by strictly verifying the communication link, ensuring the reliable transmission of robot control instructions.
[0040] Once the hardware self-test passes, parameters are initialized correctly, and the communication link is successfully established, the system automatically transitions to the enabled state and begins preparations for subsequent operation. If any item is found not to meet the requirements during the self-test, the system enters the fault state, making it easier for operators to troubleshoot and resolve the problem. This clear state transition mechanism ensures the orderliness and safety of the robot's operational process. The robot enters the enabled state only when all initialization conditions are met, avoiding the risks associated with starting the robot in the presence of hardware failures, parameter errors, or communication issues. For example, entering the enabled state before the hardware self-test fails may cause the robot to operate in a faulty state, further damaging the hardware or causing a safety incident. However, once the robot enters the fault state, relevant instructions or logs can help technicians quickly locate and resolve the problem.
[0041] Through the above-mentioned initialization state steps, the present invention can perform comprehensive and strict inspection and configuration of hardware, parameters and communication links during the robot startup phase, greatly improving the reliability and stability of the robot system startup, effectively reducing the failure rate caused by improper initialization, and providing a solid basic guarantee for the subsequent normal operation of the robot.
[0042] In one embodiment of the present invention, the enabled state of the state machine includes: The soft start process uses a progressive torque loading strategy to increase the initial torque from low to high according to the preset time interval until full load is reached; Communication quality verification: real-time monitoring of communication packet loss rate. If the packet loss rate exceeds the threshold alarm value within a preset period, a fault state is triggered; Activate the watchdog timer and set a detection period. If no heartbeat signal is received within the detection period, perform an emergency shutdown. The state transition condition is enabled, and the state automatically jumps to the motor control state after the soft start is completed and the communication quality is stable.
[0043] As described in the above steps, the present invention realizes the smooth start-up of the motor by performing progressive torque loading, real-time monitoring of the communication packet loss rate, setting the heartbeat signal detection cycle and other steps in the enabled state, while ensuring the stability of communication quality, timely discovering and handling abnormal situations, and finally enabling the robot to enter the motor control state safely and reliably, providing guarantee for the normal operation of the robot.
[0044] After the robot system completes initialization, it enters the enabled state, at which point the motors begin operating. If the motors are suddenly loaded to full torque at startup, this can significantly impact the mechanical structure, shortening the equipment's lifespan or even causing mechanical damage. The stability of the communication link is crucial for the robot to accurately receive and execute control commands. A high packet loss rate during communication can result in incomplete or inaccurate command transmission, leading to abnormal robot movement. Furthermore, if the system cannot promptly detect the operating status of the motors or other key components (for example, by using heartbeat signals to determine whether the equipment is operating properly), it may be unable to respond promptly when a fault occurs, leading to more serious problems. Therefore, a series of measures must be taken in the enabled state to address these issues.
[0045] When enabled, traditional robots use relatively crude motor startup methods, such as starting directly at full load torque, which can easily impact the mechanical structure. Monitoring of communication packet loss rates is not real-time and accurate enough, often detecting communication problems only after significant motion anomalies occur. Heartbeat signal detection mechanisms can be inadequate, with fixed and irrational detection cycles, making it difficult to detect equipment failures or misjudgments. This invention uses a progressive torque loading strategy to ensure smooth motor startup. By accurately monitoring the communication packet loss rate in real time, communication anomalies can be detected and addressed promptly. A reasonable and adjustable heartbeat signal detection cycle accurately determines the operating status of the device, comprehensively addressing the problems of traditional solutions.
[0046] Specifically, a progressive torque loading strategy is adopted, increasing the initial torque from low to high at preset time intervals. For example, the initial torque can be set to 10% of the rated torque, and then the torque is increased by 5% every 50 milliseconds until full load is reached. The motor torque is precisely controlled and adjusted by the motor drive node, during which the motor torque data is collected and fed back in real time by the motor drive node. This progressive torque loading method can effectively avoid the impact of the instantaneous high torque during motor startup on the mechanical structure. Taking an industrial robot arm as an example, if it is started directly with full-load torque, the joints and connectors of the robot arm may wear, loosen, or even be damaged due to the excessive impact force, affecting the robot's precision and service life. However, through progressive loading, the mechanical structure can gradually adapt to the torque changes and smoothly enter the working state, extending the service life of the equipment while also improving the smoothness and accuracy of the robot's movement.
[0047] During communication between the master and slave nodes, the packet loss rate is calculated by counting the number of data frames sent and received in real time. Specifically, the Ethernet control automation technology bus communication module monitors and counts the transmitted data. When the packet loss rate exceeds a threshold alarm value (such as 0.5%, which can be adjusted based on actual conditions) within a preset period (such as 100 milliseconds), a fault state is immediately triggered. Communication packet loss prevents control instructions from being fully and accurately transmitted to the slave motor drive node, causing deviations in robot motion. For example, when a robot is performing precision assembly work, packet loss can cause deviations in the robot arm's trajectory, making it impossible to accurately grasp or place parts, affecting work quality or even causing the operation to fail. By monitoring the packet loss rate in real time and triggering a fault state promptly, communication problems can be quickly discovered and resolved, ensuring the reliable transmission and execution of robot control instructions and improving the stability and reliability of robot operation.
[0048] By setting a detection cycle, the master node periodically sends heartbeat signals to slave nodes and waits for responses from them. For example, the detection cycle can be set to 500 milliseconds. If no heartbeat signal is received within this detection period, an emergency shutdown is executed. Heartbeat signals are sent and received via the Ethernet control automation technology bus, and the signal status is monitored in real time by the master node. Heartbeat signals are crucial for determining the proper functioning of slave devices (such as motor drive nodes). If a slave device malfunctions or communication is interrupted, the heartbeat signal may not be sent in a timely manner. For example, if a motor drive node stops operating due to hardware failure or software error during robot operation and fails to return a heartbeat signal, a timely emergency shutdown can be initiated through the configured detection mechanism. This prevents continued operation of the robot despite a partial device failure, potentially leading to more serious safety incidents or equipment damage, thereby ensuring the safe and reliable operation of the robot system.
[0049] When the soft start is complete (i.e., the motor torque reaches full load and communication quality is stable (packet loss rate is within the normal range and there are no heartbeat signal anomalies), the system automatically transitions to the motor control state and begins executing the robot's specific motion control tasks. A clear state transition mechanism ensures the orderly operation of the robot's operation process. The robot enters the motor control state only when the motor starts smoothly and communication is stable, avoiding the failures and risks caused by entering the operation state during startup anomalies or communication issues. For example, if the motor does not start smoothly or communication is unstable before entering the motor control state, the robot may experience abnormal motion or control failure. Strict state transition conditions effectively ensure the safety and reliability of the robot's subsequent operation.
[0050] Through the above-mentioned enabling state steps, the present invention can enable the motor to start smoothly and reduce the impact on the mechanical structure; accurately monitor the communication status in real time, promptly discover and handle communication anomalies; accurately judge the equipment operation status, and avoid safety accidents caused by untimely discovery of equipment failures, thereby providing reliable guarantees for the robot to enter the motor control state and operate normally, and improving the overall stability and safety of the robot system.
[0051] In one embodiment of the present invention, the steps of implementing the three-layer protection mechanism include: Dynamically adjust the limit range according to the real-time movement speed of the joint; Read the current joint speed value and the maximum allowable speed value preset by the system; Calculating a speed scaling factor, where the speed scaling factor is equal to the absolute value of the current speed divided by the maximum allowed speed; Calculating a dynamic limit range, where the value of the dynamic limit range is the basic limit range multiplied by a correction coefficient, for example, the correction coefficient is equal to 1 minus 0.05 times the speed proportional factor; When it is detected that the joint position exceeds the dynamic limit range, a joint over-limit exception is triggered and damping control is performed; Set the proportional gain to zero and apply a high differential gain, for example, the differential gain value must be greater than 2 times the square root of the product of the joint inertia and the stiffness coefficient.
[0052] As described in the above steps, the present invention implements the software layer protection related steps in the three-layer protection mechanism to dynamically adjust the limit range according to the real-time movement speed of the joint, and triggers damping control when the joint position exceeds the dynamic limit range, thereby effectively preventing the robot joint movement from exceeding the limit, avoiding damage to the mechanical structure, and ensuring the safe and stable operation of the robot.
[0053] During robot motion, the speed of joint movement may vary. If fixed limits are used, when the joint moves at high speed, due to factors such as inertia, even if braking is performed before reaching the limit position, the braking distance may be too long and the limit may be exceeded, causing damage to the mechanical structure. Conversely, when the joint moves at low speed, the fixed limits may restrict its normal range of motion, reducing the robot's operating efficiency. Therefore, it is necessary to dynamically adjust the limit range according to the joint movement speed and implement effective control measures, such as damping control, when the limit is exceeded to avoid these problems.
[0054] Traditional robot limit control often uses fixed limit ranges, which cannot adapt to changes in joint speed. This makes it easy to over-limit at high speeds and restricts movement flexibility at low speeds. Furthermore, when over-limit occurs, there is a lack of effective braking control strategies, often relying on mechanical hard limits to block it, which can easily cause large impacts and damage mechanical components. The present invention obtains joint speed in real time and dynamically adjusts the limit range based on a speed proportional factor, making the limit range adaptive to joint movement speed. When a joint over-limits, a damping control strategy is adopted to change the control parameters for rapid braking, preventing damage to the mechanical structure, effectively solving the problems of traditional solutions.
[0055] Specifically, speed sensors (such as encoders) installed at the joints collect the current joint speed in real time. This speed data is transmitted via a communication link to the software-layer protection unit for processing. The system's preset maximum allowable speed is stored in software configuration parameters and can be accessed at any time. Accurately acquiring the current joint speed is essential for implementing dynamic limit limits. Only by knowing the actual joint speed can the limit range be appropriately adjusted based on the speed. For example, when the robot is running fast, the joint speed is high, requiring stricter limit limits to prevent over-limiting. However, when moving slowly, the limit range can be appropriately relaxed to improve movement flexibility. The preset maximum allowable speed value provides a reference for calculating the speed scaling factor, which measures the relative speed of the current speed. The speed scaling factor is calculated by dividing the current absolute speed by the maximum allowable speed using a software algorithm. For example, if the current absolute joint speed is 50° / s and the maximum allowable speed is 100° / s, the speed scaling factor is 0.5. The speed scaling factor reflects the ratio of the current joint speed to the maximum allowable speed and is a key parameter in the subsequent calculation of the dynamic limit range. The speed scale factor quantifies the impact of the current speed on limit range adjustment, providing a basis for dynamic adjustment. The dynamic limit range is calculated by multiplying the base limit range by a correction factor. For example, the correction factor is 1 minus 0.05 times the speed scale factor. For a base limit range of ±90° and a speed scale factor of 0.5, the correction factor is 1-0.05×0.5=0.975, resulting in a dynamic limit range of ±90°×0.975=±87.75°. The dynamic limit range calculated based on the speed scale factor can be adaptively adjusted based on joint speed. Faster speeds result in smaller correction factors and narrower dynamic limit ranges, effectively preventing over-limits during high-speed motion. Slower speeds result in wider dynamic limit ranges, ensuring ample joint movement and improving robot safety and flexibility.
[0056] When the software layer protection unit detects that the joint position exceeds the dynamic limit range, it immediately triggers a joint over-limit abnormality signal and switches the control mode to perform damping control. Damping control is achieved by setting the proportional gain to zero and applying a high differential gain (the differential gain value must be greater than 2 times the square root of the product of the joint inertia and the stiffness coefficient). Specifically, the motor control parameters are adjusted through the motor drive node. By promptly detecting the joint over-limit abnormality and performing damping control, the movement speed of the joint can be quickly reduced when it exceeds the safety range, avoiding rigid collisions with the mechanical hard limit and reducing damage to the mechanical structure caused by the impact force. For example, when the robot arm moves rapidly and approaches the limit position, the damping control can take effect at the moment of over-limit, causing the arm to quickly slow down and stop. While protecting the mechanical structure, it can also prevent problems such as robot movement out of control due to collisions.
[0057] Through the above steps, the present invention realizes the dynamic limit and damping control functions in the software layer protection, making the robot joint limit control more intelligent and flexible, effectively avoiding the risk of mechanical damage caused by joint over-limit, improving the safety and reliability of the robot operation, and also optimizing the robot's motion performance at different speeds.
[0058] In one embodiment of the present invention, the steps of implementing the three-layer protection mechanism include: Adopt sliding window detection method to store multiple frames of continuous motor angle data; When the angle difference between two adjacent frames exceeds the preset arc, the second-order verification is started to check whether the direction of the torque change rate matches the motion trend; If the data is abnormal within the preset frame period, an emergency stop is triggered; Set basic timeout thresholds and calculate network jitter coefficients; Dynamically generate a timeout threshold based on the basic timeout threshold and the calculated network jitter coefficient; When the instruction response time exceeds the actual timeout threshold, a protection action is triggered; Real-time monitoring of the current motor temperature and ambient temperature; When the temperature change rate is greater than the heat dissipation coefficient multiplied by the difference between the current temperature and the ambient temperature, an overheating warning is triggered.
[0059] As described in the above steps, the present invention adopts a series of measures such as sliding window detection method, second-order verification, dynamic generation of timeout thresholds, and real-time monitoring of temperature change rate in software layer protection to achieve anomaly detection of multi-dimensional data such as motor angle, torque, command response time and temperature during the operation of the robot, timely discover and handle potential faults, and ensure the safe and stable operation of the robot.
[0060] During robot operation, unusual changes in parameters such as motor angle, torque, command response time, and temperature can indicate potential faults. For example, a sudden change in motor angle could indicate a sensor failure or mechanical structure anomaly; a mismatch between the torque rate of change and the motion trend could indicate an abnormal load or control algorithm error; a long command response time could indicate a communication failure or processor performance degradation; and unusual temperature changes could indicate motor heat dissipation issues or overload. Failure to detect and address these anomalies in a timely manner could result in robot damage or even safety incidents, necessitating comprehensive, real-time monitoring and anomaly detection of these parameters.
[0061] Traditional robot anomaly detection mostly uses single-dimensional data detection, such as monitoring only motor current or temperature, which cannot comprehensively judge complex faults; the detection method is simple and lacks effective filtering means for interference such as sensor noise, which is prone to false alarms; the timeout threshold is fixed and cannot adapt to dynamic changes such as network jitter, which may lead to false triggering of protection actions under normal circumstances or failure to respond in time when a fault occurs. This invention adopts multi-dimensional data fusion detection, combined with advanced algorithms such as sliding window detection and second-order verification, to comprehensively monitor motor angle, torque, command response time, and temperature; dynamically generate timeout thresholds to adapt to network jitter; and use heat dissipation equations to detect temperature anomalies, effectively solving the problems existing in traditional solutions.
[0062] Specifically, a sliding window detection method is used. The software-layer protection unit continuously stores multiple frames of motor angle data (e.g., a window of five frames, which is updated over time). When the angle difference between two adjacent frames exceeds a preset value (e.g., 0.6 radians, adjustable based on actual conditions), a second-order verification is initiated to check whether the direction of the torque rate of change matches the motion trend. Torque data is collected in real time by a torque sensor installed on the motor and transmitted to the software-layer protection unit. The sliding window detection method captures continuous changes in motor angle and promptly detects sudden changes in angle. The second-order verification effectively filters out interference factors such as sensor noise, preventing false alarms. For example, if the angle data jumps solely due to transient sensor noise, by checking the direction of the torque rate of change, if the torque does not change synchronously, the data can be considered invalid and the shutdown protection will not be triggered, ensuring normal robot operation. However, if the sudden angle change is caused by a mechanical structural abnormality or other factors, and the torque rate of change direction does not match, the problem can be detected and an emergency shutdown can be triggered to prevent the fault from escalating.
[0063] First, set a basic timeout threshold (for example, 500 milliseconds, adjustable based on system requirements). Then, by monitoring data latency and jitter during network transmission, calculate the network jitter coefficient (for example, based on statistics such as the variance of transmission delay over a certain period). Dynamically generate a timeout threshold based on the basic timeout threshold and the calculated network jitter coefficient. The formula can be simply expressed as: Actual timeout threshold = Basic timeout threshold × (1 + Network jitter coefficient). When the command response time exceeds the actual timeout threshold, protection is triggered. In real-world network environments, network jitter is inevitable. Fixed timeout thresholds cannot adapt to this dynamic variation, potentially leading to false triggering of protection actions (determining a fault even with a slight network delay) or inability to respond promptly to a fault (due to high network jitter, a fault may occur but not reach the fixed timeout threshold, thus preventing protection from being triggered). The dynamically generated timeout threshold can be adaptively adjusted according to the actual network jitter. When the network jitter is small, the timeout threshold is kept low to ensure timely response. When the network jitter is large, the threshold is appropriately relaxed to avoid false triggering. This ensures that the command response can be accurately judged in various network environments, ensuring the reliable transmission and execution of robot control commands.
[0064] A temperature sensor installed on the motor monitors the current motor temperature in real time, while an ambient temperature sensor monitors the ambient temperature. The software-layer protection unit calculates the temperature change rate in real time. When the temperature change rate is greater than the heat dissipation coefficient multiplied by the difference between the current temperature and the ambient temperature (the heat dissipation coefficient is predetermined based on the motor material, heat dissipation structure, etc.), an overheating warning is triggered. An abnormal increase in motor temperature may be caused by overload, poor heat dissipation, or other reasons. By monitoring the temperature change rate in real time and comparing it with a threshold calculated based on the heat dissipation equation, motor overheating trends can be detected in advance. For example, if the cooling system malfunctions during long-term high-load operation of the robot, the motor temperature will rise rapidly. When the temperature change rate exceeds the threshold, an overheating warning is triggered, prompting the operator to take timely measures (such as reducing the load, checking the cooling system, etc.) to prevent motor damage due to overheating and ensure reliable operation of the robot.
[0065] Through the above steps, the present invention realizes the multi-dimensional anomaly detection function in software layer protection, which can effectively filter interference, adapt to dynamic changes in the network, and provide early warning of temperature anomalies, greatly improving the accuracy and timeliness of robot fault detection, reducing the risk of equipment damage caused by faults, and improving the safety and reliability of robot operation.
[0066] The present invention also discloses a robot whole-body motion control system based on reinforcement learning, comprising: An inference node, configured to collect inertial measurement unit data and motor feedback data, and to input the inertial measurement unit data and the motor feedback data into a deep reinforcement learning model accelerated by TensorRT to generate motor control instructions; The master node is configured to issue the control instructions via the Ethernet control automation technology bus, wherein the upper layer of the master node runs a four-state state machine including an initialization state, an enable state, a motor control state, and a fault state; and the lower layer implements, through a protocol stack, the following: converting the control instructions into service data object data packets, implementing clock synchronization using a distributed clock synchronization mechanism, and collecting and updating slave device status words in real time; A motor driving node is used to parse the service data object data packet and drive the motor to realize the robot joint motion control; The protection module is used to implement a three-layer protection mechanism, including electrical layer protection, mechanical layer protection and software layer protection, which are respectively realized through the electrical layer protection unit, mechanical layer protection unit and software layer protection unit. The electrical layer protection unit is used to integrate the dynamic heartbeat detection circuit and the overcurrent protection relay; the mechanical layer protection unit is used to install the hard limit switch and the buffer zone detection sensor; the software layer protection unit is used to deploy the dynamic soft limit algorithm and the multi-dimensional anomaly detection program.
[0067] As described in the above modules, the present invention achieves a systematic improvement in the real-time, scalability and safety of the robot's whole-body motion control by constructing a layered hardware architecture of "inference node-master node-motor drive node-protection module" and combining it with a deep reinforcement learning model, Ethernet control automation technology bus communication and a multi-layer protection mechanism, thereby solving the problems of tight module coupling, single protection system and insufficient generalization in traditional control systems.
[0068] Because motion control requires both intelligent decision-making and real-time execution, traditional integrated architectures (such as single-chip microcontrollers directly controlling motors) struggle to simultaneously meet the computational power requirements of DRL model inference and the microsecond-level instruction issuance requirements. Multi-joint robots (such as 30-joint humanoid robots) require the coordinated operation of multiple modules. High inter-module coupling (e.g., a hybrid design of control logic and communication protocols) results in poor scalability (e.g., adding new sensors requires rewriting the entire code). Physical risks such as electrical overcurrent and mechanical collision require dedicated hardware circuits (e.g., MOSFET drivers and hard limit switches) for rapid response. Relying solely on software algorithms (e.g., pure software overcurrent detection) can lead to protection failures due to system latency. This invention, through hardware layer decoupling, communication protocol hardening, and physical redundancy of protection mechanisms, constructs a robot control system that combines intelligent decision-making, real-time control, and multi-layered safety protection. Its core breakthrough lies in the deep integration of algorithmic intelligence and hardware real-time performance, resolving the conflict between high computational power requirements and low-latency control in traditional systems. This provides a scalable hardware architecture solution for complex scenarios such as industrial robots and service robots.
[0069] The master station node includes: Initialization status unit, performs hardware self-test and parameter initialization, and activates the fault state when a mismatch in the number of slaves is detected; Enable state unit, used to output torque according to the gradient increasing strategy; and used to calculate the packet loss rate in real time and control the watchdog timer; The motor control state unit is used to support the expansion of motion sub-states, including independent control logic for walking, standing, and jumping; and is used to plan joint motion curves based on deep reinforcement learning output instructions; The fault status unit is used to distinguish between hardware faults and software faults. Hardware faults include overcurrent and overtemperature, and software faults include communication timeout and data anomalies; and is used to execute joint position locking or system safety shutdown.
[0070] As described in the above units, the present invention realizes the standardization, scalability and precise fault handling of the robot control process by setting the initialization state unit, the enable state unit, the motor control state unit and the fault state unit in the state machine management module, ensures the atomicity and safety of operations in each operation stage, and solves the problems of the traditional state machine's lack of sub-state expansion capability and fuzzy fault classification.
[0071] Since the process from starting up to executing motion requires the robot to go through the stages of "hardware self-test → safe start → motion control → exception handling", each stage must be executed in strict sequence to avoid hardware damage due to process confusion (for example, starting the motor directly without self-test may cause overcurrent). Multiple motion modes (such as walking, standing, and jumping) require independent control logic. Traditional fixed state machines are difficult to expand dynamically, resulting in code redundancy (for example, the state machine logic needs to be rewritten to add a "climbing" mode). The handling of hardware failures (such as motor overcurrent) and software failures (such as communication timeouts) is significantly different. The traditional "one-size-fits-all" shutdown may lead to misoperation (such as an emergency shutdown triggered by a momentary network fluctuation). The motor control state unit of the present invention supports sub-state expansion (such as independent logic for walking / standing / jumping) and realizes mode switching by loading different motion sub-state parameters without modifying the main state machine code. The fault state unit distinguishes hardware / software faults through an error classifier. Hardware faults directly trigger hardware shutdown, and software faults prioritize redundant communication to improve the system's fault tolerance. The enabling state unit uses gradient torque loading and communication quality verification to avoid control failure caused by motor startup impact and invalid communication.
[0072] The software layer protection unit includes: A dynamic limit calculator, used to receive joint velocity data in real time, calculate velocity scaling factors, and output dynamic limit thresholds; Damping controller, used to switch control mode when over-limit is triggered; The anomaly detector is used to store the motor angle data, compare the logical relationship between the angle change and the torque change rate, and generate an overheating warning signal based on the heat dissipation equation.
[0073] As described in the above-mentioned software layer protection unit, the present invention realizes dynamic adjustment of the robot joint motion range, rapid braking when exceeding the limit, and anomaly detection of multi-dimensional data such as motor angle, torque, and temperature by setting a dynamic limit calculator, a damping controller and anomaly detector in the software layer protection module. In conjunction with the electrical layer and mechanical layer protection, a multi-level active safety protection system is formed to ensure the safe and stable operation of the robot.
[0074] At high speeds, inertia can cause a robot joint to exceed its fixed limits, leading to collision and damage. At low speeds, fixed limits can restrict its flexibility. Therefore, it's necessary to dynamically adjust the limits based on the joint's real-time speed to balance safety and efficiency.
[0075] Traditional control methods have slow braking response when joints exceed limits, which can easily cause mechanical shock. However, fast and effective damping control can reduce shock and protect mechanical structures. Furthermore, single sensor data can be affected by noise, leading to false alarms or missed detections. This invention combines multi-dimensional data verification, including torque, temperature, and network, to more accurately determine true anomalies and avoid misoperation.
[0076] Specifically, the real-time joint speed is obtained through the motor encoder, compared with the system's preset maximum allowable speed, and a speed scaling factor is calculated. The limit range is then dynamically adjusted based on the formula, and the result is sent to the motor drive node in real time. For example, when the robot is running, the joint speed is fast, and the limit range is automatically narrowed to prevent overtravel due to inertia. When performing delicate operations, the speed is slow, and the limit range is relaxed to ensure flexible movement.
[0077] When a joint position exceeds a dynamic limit or approaches a hard limit warning zone, the proportional gain is immediately set to zero, a high differential gain is applied and decays over time, and the motor torque is adjusted through the motor drive node to rapidly decelerate the joint. Compared to traditional braking methods, this solution can reduce joint speed in a shorter time and reduce impact energy. For example, if a robotic arm exceeds a limit, it can stop quickly and smoothly, avoiding collision damage.
[0078] The sliding window processor stores the most recent frames of motor angle data and calculates the angle difference between adjacent frames. When the angle difference exceeds a threshold, it combines it with torque data to verify consistent motion trends and eliminate sensor noise interference. The processor also monitors the motor temperature change rate in real time and compares it with a preset heat dissipation model to provide early warning of overheating risks. For example, if the angle data in a frame suddenly changes but the torque remains unchanged, it is identified as noise rather than a true anomaly, avoiding premature shutdowns. If the temperature changes abnormally, it prompts cooling or load reduction to prevent motor overheating and damage.
[0079] In summary, dynamic limiting achieves speed adaptive protection and reduces the risk of high-speed overtravel; damping control significantly shortens braking time and reduces mechanical impact; anomaly detection can effectively filter interference, improve fault identification accuracy, significantly reduce false alarm rate, and ensure safe and reliable operation of the robot.
[0080] The present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of a robot whole-body motion control method based on reinforcement learning.
[0081] The present invention also discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the steps of a robot whole-body motion control method based on reinforcement learning.
[0082] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.
[0083] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A robot whole-body motion control method based on reinforcement learning, characterized in that: The following steps are involved: Real-time collection of inertial measurement unit data and motor feedback data through inference nodes; Inputting the inertial measurement unit data and the motor feedback data into a deep reinforcement learning model accelerated by TensorRT to generate motor control instructions; The control instructions are issued by the master node using the Ethernet control automation technology bus, wherein the upper layer of the master node runs a four-state state machine including an initialization state, an enable state, a motor control state, and a fault state; the lower layer implements, through a protocol stack, the following: converting the control instructions into service data object data packets, using a distributed clock synchronization mechanism to achieve clock synchronization, and real-time acquisition and updating of slave device status words; The motor drive node parses the service data object data packet and drives the motor to realize the robot joint motion control; A three-layer protection mechanism is implemented during the above steps, including electrical layer protection, mechanical layer protection and software layer protection.
2. A robot whole body motion control method based on reinforcement learning according to claim 1, characterized in that: The initialization state of the state machine includes: Perform hardware self-test, check emergency stop signal and Ethernet control automation technology bus ring network connection status, and verify the number of slave stations and configuration matching; Initialize parameters, load software configuration parameters and control parameters, including joint motion range threshold and current protection threshold; Establishing a communication link, completing the handshake protocol between the Ethernet control automation technology bus master station and the slave station, and determining that the connection is successful when the packet loss rate is less than a preset threshold; Initialize the state transition condition, automatically jump to the enabled state after the self-test passes, and enter the fault state if the self-test fails.
3. The robot whole-body motion control method based on reinforcement learning according to claim 1, characterized in that: The enabled states of the state machine include: The soft start process uses a progressive torque loading strategy to increase the initial torque from low to high according to the preset time interval until full load is reached; Communication quality verification: real-time monitoring of communication packet loss rate. If the packet loss rate exceeds the threshold alarm value within a preset period, a fault state is triggered; Activate the watchdog timer and set a detection period. If no heartbeat signal is received within the detection period, perform an emergency shutdown. The state transition condition is enabled, and the state automatically jumps to the motor control state after the soft start is completed and the communication quality is stable.
4. The robot whole-body motion control method based on reinforcement learning according to claim 1, characterized in that: The steps for implementing the three-layer protection mechanism include: Dynamically adjust the limit range according to the real-time movement speed of the joint; When it is detected that the joint position exceeds the dynamic limit range, a joint over-limit exception is triggered and damping control is performed; Set the proportional gain to zero and apply a high derivative gain.
5. The robot whole body motion control method based on reinforcement learning according to claim 4 is characterized in that: The steps for implementing the three-layer protection mechanism include: Adopt sliding window detection method to store multiple frames of continuous motor angle data; When the angle difference between two adjacent frames exceeds the preset arc, the second-order verification is started to check whether the direction of the torque change rate matches the motion trend; If the data is abnormal within the preset frame period, an emergency stop is triggered; Set basic timeout thresholds and calculate network jitter coefficients; Dynamically generate a timeout threshold based on the basic timeout threshold and the calculated network jitter coefficient; When the instruction response time exceeds the actual timeout threshold, a protection action is triggered; Real-time monitoring of the current motor temperature and ambient temperature; When the temperature change rate is greater than the heat dissipation coefficient multiplied by the difference between the current temperature and the ambient temperature, an overheating warning is triggered.
6. A robot whole-body motion control system based on reinforcement learning, characterized in that: include: An inference node, configured to collect inertial measurement unit data and motor feedback data, and to input the inertial measurement unit data and the motor feedback data into a deep reinforcement learning model accelerated by TensorRT to generate motor control instructions; The master node is configured to issue the control instructions via the Ethernet control automation technology bus, wherein the upper layer of the master node runs a four-state state machine including an initialization state, an enable state, a motor control state, and a fault state; and the lower layer implements, through a protocol stack, the following: converting the control instructions into service data object data packets, implementing clock synchronization using a distributed clock synchronization mechanism, and collecting and updating slave device status words in real time; A motor driving node is used to parse the service data object data packet and drive the motor to realize the robot joint motion control; The protection module is used to implement a three-layer protection mechanism, including electrical layer protection, mechanical layer protection and software layer protection, which are respectively realized through the electrical layer protection unit, mechanical layer protection unit and software layer protection unit. The electrical layer protection unit is used to integrate the dynamic heartbeat detection circuit and the overcurrent protection relay; the mechanical layer protection unit is used to install the hard limit switch and the buffer zone detection sensor; the software layer protection unit is used to deploy the dynamic soft limit algorithm and the multi-dimensional anomaly detection program.
7. The robot whole-body motion control system based on reinforcement learning according to claim 6, characterized in that: The master station node includes: Initialization status unit, performs hardware self-test and parameter initialization, and activates the fault state when a mismatch in the number of slaves is detected; Enable state unit, used to output torque according to the gradient increasing strategy; and used to calculate the packet loss rate in real time and control the watchdog timer; The motor control state unit is used to support the expansion of motion sub-states, including independent control logic for walking, standing, and jumping; and is used to plan joint motion curves based on deep reinforcement learning output instructions; The fault status unit is used to distinguish between hardware faults and software faults. Hardware faults include overcurrent and overtemperature, and software faults include communication timeout and data anomalies; and is used to execute joint position locking or system safety shutdown.
8. The robot whole-body motion control system based on reinforcement learning according to claim 6, characterized in that: The software layer protection unit includes: A dynamic limit calculator, used to receive joint velocity data in real time, calculate velocity scaling factors, and output dynamic limit thresholds; Damping controller, used to switch control mode when over-limit is triggered; The anomaly detector is used to store the motor angle data, compare the logical relationship between the angle change and the torque change rate, and generate an overheating warning signal based on the heat dissipation equation.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Cited By
Fault diagnosis method, electronic equipment, storage medium and program product
CN122185247A
Fault diagnosis method, electronic device, storage medium, and program product
CN122185247B