Joint control method and system for multi-degree-of-freedom robot
By using multi-source sensor data fusion and reinforcement learning to optimize control, the problems of error accumulation and nonlinear coupling in the joint control of traditional multi-degree-of-freedom robots are solved, achieving high-precision and stable joint control and enhancing the robot's adaptive capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI DIBICO PRECISION ELECTRONICS CO LTD
- Filing Date
- 2026-05-25
- Publication Date
- 2026-06-26
AI Technical Summary
Traditional multi-degree-of-freedom robot joint control relies on a single sensor, which is prone to error accumulation, has weak dynamic adaptability, and makes it difficult to achieve high-precision control due to nonlinear coupling of joints.
Multi-source sensor data acquisition is adopted, combining accelerometer, angular velocity sensor and motor encoder, and integrating vision sensor for joint state estimation. The control strategy is optimized through fuzzy control analysis and reinforcement learning to generate joint control commands.
It improves the accuracy of joint motion positioning and operational stability, enhances the robot's adaptability under complex working conditions, and ensures reliable completion of tasks.
Smart Images

Figure CN122275012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial robot control technology, specifically to a joint control method and system for multi-degree-of-freedom robots. Background Technology
[0002] With the rapid development of intelligent manufacturing, multi-degree-of-freedom robots have been widely used in complex work scenarios such as precision assembly, logistics handling, and medical surgery. The precision of their joint control directly determines the robot's task execution capability and operational reliability. As the complexity of work tasks increases, higher requirements are placed on the real-time performance, robustness, and dynamic adaptability of robot joint control.
[0003] Currently, joint control in multi-degree-of-freedom industrial robots mostly employs PID control methods based on single motor encoder feedback. This approach suffers from large cumulative errors and the inability to detect joint flexibility deformation and external impact deviations. While some solutions utilize inertial sensor fusion, they still cannot eliminate global drift. Furthermore, traditional control methods based on precise mathematical models struggle to adapt to the strong coupling and nonlinear characteristics between joints, and pre-programmed control strategies cannot adapt to dynamic factors such as environmental changes and robot wear. This results in a significant decrease in control accuracy and task completion rate in complex environments. Summary of the Invention
[0004] This application provides a joint control method and system for multi-degree-of-freedom robots, aiming to solve the technical problems of traditional multi-degree-of-freedom robot joint control relying on a single sensor, which easily leads to error accumulation, weak dynamic adaptability, and difficulty in achieving high-precision control due to nonlinear coupling of joints.
[0005] In view of the above problems, this application provides a joint control method and system for multi-degree-of-freedom robots.
[0006] The first aspect disclosed in this application provides a joint control method for a multi-degree-of-freedom robot. The method includes: installing a sensor group on M joints of a target multi-degree-of-freedom robot, the sensor group integrating an accelerometer, an angular velocity sensor, and a motor encoder; simultaneously setting a vision sensor within the workspace of the target multi-degree-of-freedom robot; collecting M joint operation data through the sensor group and acquiring a global image of the robot collected by the vision sensor; performing joint state estimation based on the M joint operation data and the global image of the robot to obtain M joint position state information; combining the robot's working environment map and task requirements to perform motion path decomposition and fuzzy control analysis on the M joint position state information to obtain M joint control parameters; introducing a reinforcement learning mechanism to optimize the control strategy of the target multi-degree-of-freedom robot based on the M joint control parameters, determining M joint control commands, and controlling the robot's operation through the M joint control commands.
[0007] Another aspect of this application discloses a joint control system for a multi-degree-of-freedom robot. The system includes: a sensor group mounting module for mounting sensor groups on M joints of a target multi-degree-of-freedom robot, the sensor groups integrating accelerometers, angular velocity sensors, and motor encoders, and a vision sensor positioned within the workspace of the target multi-degree-of-freedom robot; a joint position information acquisition module for collecting M joint operation data through the sensor groups and acquiring a global image of the robot collected by the vision sensor, estimating joint states based on the M joint operation data and the global image to obtain M joint position state information; a joint control parameter acquisition module for performing motion path decomposition and fuzzy control analysis on the M joint position state information in conjunction with a robot work environment map and task requirements to obtain M joint control parameters; and a joint control command determination module for introducing a reinforcement learning mechanism to optimize the control strategy of the target multi-degree-of-freedom robot based on the M joint control parameters, determining M joint control commands, and controlling the robot's operation through the M joint control commands.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0009] By employing a technical solution that combines multi-source sensor data acquisition, dual-path joint position estimation, and data fusion in a base coordinate system with path planning based on the working environment map and task requirements, and obtains control parameters through fuzzy control analysis, and then introduces reinforcement learning to optimize the control strategy and output commands to control the robot, this solution solves the technical problems of traditional robot control that rely on a single sensor and are prone to error accumulation, nonlinear joint coupling, and insufficient control accuracy and adaptability. It achieves the technical effects of improving joint motion positioning accuracy and operational stability, enhancing the robot's adaptability in complex working conditions, and ensuring reliable completion of tasks.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0011] Figure 1 A flowchart illustrating a joint control method for a multi-degree-of-freedom robot is provided for embodiments of this application.
[0012] Figure 2 A schematic diagram of the joint control system of a multi-degree-of-freedom robot is provided for the embodiments of this application.
[0013] Explanation of reference numerals in the attached drawings: Sensor assembly module 11, Joint position information acquisition module 12, Joint control parameter acquisition module 13, Joint control command determination module 14. Detailed Implementation
[0014] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0015] The overall concept of the technical solution provided in this application is as follows: This application provides a joint control method and system for a multi-degree-of-freedom robot. First, multi-source sensor data is collected from the robot, and local and global joint position estimations are performed separately. The two types of data are then fused in a base coordinate system to obtain a precise joint position state. Next, a path is planned based on the environmental map and task requirements, and control parameters are generated through fuzzy control analysis. Finally, reinforcement learning is used to optimize the control commands, achieving high-precision adaptive operation control of the robot joints.
[0016] After introducing the basic principles of this application, various non-limiting embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0017] Example 1, as Figure 1 As shown in the embodiment of this application, a joint control method for a multi-degree-of-freedom robot is provided, the method comprising: Step S100: Install sensor groups on M joints of the target multi-degree-of-freedom robot. The sensor groups integrate accelerometers, angular velocity sensors and motor encoders. At the same time, set up vision sensors in the workspace of the target multi-degree-of-freedom robot.
[0018] Specifically, a target multi-degree-of-freedom robot refers to a robot with three or more independently movable joints that requires joint control. M joints refer to the total number of movable connected parts on the robot capable of relative motion, where M is a positive integer. Accelerometers are sensors used to measure the linear acceleration of an object, detecting the linear acceleration of joints in the X, Y, and Z directions. Angular velocity sensors are sensors used to measure the angular velocity of an object, detecting the rotational angular velocity of joints around the X, Y, and Z axes. Motor encoders are position sensors mounted on the joint drive motor shafts, measuring the motor's rotational angle and speed. The workspace refers to the set of all spatial points accessible to the robot's end effector; for example, the workspace of a 6-axis industrial robot with a 10kg load capacity and a 1.5-meter reach is an approximately spherical area with a radius of 1.5 meters. Vision sensors are sensors used to acquire image information of the environment or target, including industrial area scan cameras and depth cameras.
[0019] Specifically, the positions of all M independently movable joints are determined based on the specific configuration of the target multi-degree-of-freedom robot. For example, for a 6-DOF industrial robotic arm used for precision assembly of automotive parts, M=6, corresponding to the base rotation joint, upper arm pitch joint, forearm pitch joint, wrist rotation joint, wrist pitch joint, and wrist yaw joint, respectively. A pre-integrated sensor group is installed at the output end of each joint (instead of the traditional motor end). The accelerometer and angular velocity sensor are usually integrated into a small IMU module with a volume of less than 1 cm³, which is used to directly measure the actual motion state of the joint output end. The motor encoder is installed on the servo motor shaft of the corresponding joint to measure the rotational position and speed of the motor. At the same time, vision sensors are reasonably arranged in the robot's workspace, and their focal length and angle are adjusted to ensure that the field of view can completely cover the entire workspace of the robot, and that the characteristic markings of all joints of the robot can be clearly captured by the vision sensors in any working posture. Microsecond-level time synchronization calibration and spatial coordinate calibration are performed on all sensors to establish a unified time reference and accurate spatial coordinate transformation relationship between the sensors, laying the foundation for subsequent multi-source data fusion processing.
[0020] Preferably, the inertial measurement unit (IMU) module, which integrates accelerometer and angular velocity sensors, must be installed at the joint output end, not the motor end. Specifically, it should be fixed to the output flange of the joint link or the link body near the joint axis of rotation, and must not be installed on the servo motor housing, the reducer input end, or the transmission mechanism. For example, for the pitch joint of a multi-degree-of-freedom industrial robotic arm, the IMU module should be installed on the arm link near the joint rotation axis, rather than on the motor housing that drives the arm.
[0021] The straight-line distance between the IMU module's mounting point and the corresponding joint's rotation axis should be less than 50mm to minimize the interference of centrifugal acceleration and Coriolis acceleration on the measurement results. The IMU module should be avoided in areas of severe vibration, such as the reducer housing and motor heat sink, as well as areas of significant deformation under stress on the connecting rods. The motor encoder must be coaxially mounted with the joint drive motor shaft, preferably on the non-load end of the motor, i.e., the rear end cover, to facilitate installation and maintenance and avoid damage from load impacts. Each independently movable joint must have its own complete sensor set; sensors must not be shared across joints. The installation positions of all sensors must not alter the robot's original mechanical structure and range of motion, must not interfere with the movement of other robot components or peripheral equipment, and must not obscure the pre-set visual feature markings on the joints.
[0022] The installation requirements for the joint-end sensor assembly are as follows: Mechanical fixing should employ a double-fixing method using rigid threaded connections and locating pins; non-rigid connections such as adhesives and clips are prohibited. Each sensor module should use at least two M3 or larger fastening screws and one cylindrical locating pin to ensure no loosening or displacement under high-speed robot movement and impact loads. The flatness error of the sensor mounting surface should be less than 0.05mm, and the surface roughness Ra ≤ 1.6μm, ensuring a completely tight fit between the sensor and the mounting surface. The coaxiality error between the motor encoder shaft and the motor shaft should be less than 0.02mm to avoid periodic measurement errors caused by eccentricity. The sensor cable should be a shielded twisted-pair cable, routed inside the robot link, with a minimum distance of 100mm from the power cable, and should not be parallel to it. The cable connector should use an aviation plug and be waterproof and dustproof, with a protection level of at least IP65 in industrial environments and at least IP67 in medical or humid environments. An aluminum alloy shield should be added to the outside of the sensor module to suppress electromagnetic radiation interference in the industrial environment.
[0023] The installation requirements for the global vision sensor are as follows: A top-view installation is preferred, positioned 2-5 meters directly above the robot's workspace. Adjust the focal length and angle to ensure the field of view completely covers the entire workspace, and that visual feature markers of all joints are clearly captured in any working posture. If top-view installation is limited, a multi-camera surround installation can be used, with 2-4 vision sensors evenly distributed around the workspace. Global coverage is achieved through multi-view image fusion. The vision sensor should be installed away from strong light sources, reflective objects, and vibration sources to avoid image overexposure, reflection, and motion blur. The vision sensor should be fixed with a rigid metal bracket with height and angle fine-tuning capabilities for easy adjustment of the field of view.
[0024] After the articulated sensor assembly is installed, spatial attitude calibration must be performed: A high-precision coordinate measuring machine (CMM) is used to determine the transformation matrix between the IMU coordinate system and the corresponding joint coordinate system; the calibration angle error should be less than 0.1°. All sensors, including the articulated sensor assembly and the global vision sensor, must be connected to a unified hardware synchronization clock system to achieve microsecond-level time synchronization, ensuring strict time alignment of data collected by different sensors. After the global vision sensor is installed, intrinsic and extrinsic parameter calibration must be performed: A checkerboard calibration board is used to determine the camera's intrinsic parameters, including focal length, principal point coordinates, distortion coefficients, and extrinsic parameters, i.e., the camera's position and attitude in the global coordinate system; the calibration reprojection error should be less than 0.5 pixels.
[0025] This step, by integrating multiple types of sensors at the joint output end and combining them with a global vision sensor, avoids the core problems of traditional single-motor encoders, which cannot detect joint flexibility deformation and actual position deviation caused by external impacts, and suffer from cumulative errors. The integrated sensor group design significantly reduces installation space and wiring complexity, lowering the system failure rate. The hardware synchronous triggering mechanism ensures the temporal consistency of multi-source data, providing high-quality raw data support for subsequent multi-source fusion joint state estimation. At the same time, the redundant design of multiple sensors gives the system a certain degree of fault tolerance. When one sensor experiences a minor failure, other sensors can still provide basic state information, ensuring the safe operation of the robot.
[0026] Step S200: Collect running data of M joints through the sensor group, and simultaneously acquire the global image of the robot collected by the vision sensor. Based on the running data of the M joints and the global image of the robot, perform joint state estimation to obtain the position state information of the M joints.
[0027] Specifically, joint motion data refers to the raw, heterogeneous data collected by the sensor array that reflects the true motion state of the joints. This includes three-dimensional linear acceleration output from accelerometers, three-dimensional angular velocity output from angular velocity sensors, and motor rotation angles and speeds output from motor encoders. The robot global image refers to single-frame or continuous-frame images collected by vision sensors within the workspace, completely covering all joints of the robot, used to provide a global absolute position reference for the joints. Joint position state information refers to the precise three-dimensional position and attitude data of the joints obtained through multi-source fusion and unified in the robot's base coordinate system; this is the sole input for subsequent motion planning and control.
[0028] Specifically, a unified hardware synchronization clock system is built to trigger the synchronous acquisition of all sensors. The joint-end sensor group acquires the running data of M joints at a high sampling rate of 1kHz. The sensor group of each joint will synchronously output the three-dimensional linear acceleration, three-dimensional angular velocity and motor rotation angle data of that joint. At the same time, the global binocular vision sensor installed directly above the workspace acquires the global image of the robot at a frame rate of 30Hz, ensuring that each frame can clearly capture the unique ArUco mark pasted on all joints. The acquired raw data is preprocessed to obtain the running data of M available joints and the global image of the robot at the corresponding time.
[0029] Independent state estimation is performed on two types of data. On one hand, the available joint operation data is processed using M joint coordinate systems. The first joint position information is obtained by numerical integration of acceleration and angular velocity data. This information is then compared with the rotational position data from the motor encoder for error verification. If the comparison result is within a preset error threshold, the two data are fused to obtain the initial joint position information. If the threshold is exceeded, the encoder data is deemed invalid, and only the inertial integration result is used as the initial joint position information. On the other hand, the global coordinate system is used to process the robot's global image. Template matching detection is performed using pre-established M joint feature templates. The ArUco markers of each joint are identified and their pixel coordinates are extracted. Then, 3D reconstruction is performed by combining the pre-calibrated intrinsic and extrinsic parameters of the vision sensor to generate the second joint position information. This information is then converted into global joint position information in the global coordinate system. Based on the pre-calibrated robot base coordinate system, the coordinate transformation relationship matrices between the joint coordinate system, the global coordinate system, and the base coordinate system are obtained. The initial joint position information and the global joint position information are uniformly transformed into the base coordinate system. The fusion weights of the two types of data are dynamically adjusted according to the current motion state of the robot to perform weighted fusion estimation, thereby obtaining the precise position state information of M joints in the base coordinate system.
[0030] This step maximizes the complementary advantages of each sensor by combining local joint sensing with global visual positioning through a dual-path independent state estimation and dynamic weighted fusion mechanism. It retains the advantages of high sampling rate and high dynamic response of the joint-end sensors, which can accurately capture the instantaneous state of the joint during high-speed movement. At the same time, it completely eliminates the cumulative errors of single inertial measurement and encoder measurement by using the absolute position reference provided by the global visual sensor. Through the error verification mechanism of encoder and inertial data, it effectively solves the problems of encoder slippage, disconnection, electromagnetic interference, as well as the deviation between the actual position and the motor command position caused by joint flexibility deformation, external impact, and transmission backlash, which significantly improves the robustness of state estimation. By dynamically adjusting the fusion weights, it can adapt to the differences in sensor performance under different motion conditions, providing a reliable state feedback basis for subsequent high-precision motion planning and control.
[0031] Step S300: Combine the robot's working environment map and task requirements to perform motion path decomposition and fuzzy control analysis on the position and state information of the M joints to obtain the control parameters of the M joints.
[0032] Specifically, a robot work environment map refers to a pre-constructed two-dimensional or three-dimensional grid map that includes all static obstacles, equipment locations, and safety zones within the robot's workspace, serving as the spatial constraint basis for path planning. Task requirements refer to the specific operational objectives the robot needs to accomplish, including the target position, posture, speed, accuracy requirements, and sequence of operations for the end effector.
[0033] Motion path decomposition: This is the process of inversely decomposing the global task path of the robot's end effector into the independent motion trajectory of each joint, based on the robot's kinematic characteristics. Joint control parameters refer to the control quantities used to drive the joint servo motors, including the joint target angle, angular velocity, angular acceleration, and torque, which are the direct inputs to the servo controller.
[0034] Specifically, a pre-built robot working environment map and specific task requirements are loaded; then, the environment map is combined with A... The algorithm performs global path planning, generating a collision-free global task path from the robot's end effector to the screw feeder and then to the battery module. It acquires the mechanical characteristics of M joints, including maximum angular velocity, maximum angular acceleration, torque limits, and range of motion constraints, ensuring the generated motion trajectory conforms to the robot's physical limits. Using an inverse kinematics algorithm, each path point on the global task path is decomposed into the target angle of each joint at the corresponding moment, generating the planned motion trajectory for M joints. A joint fuzzy controller is constructed, selecting joint trajectory deviation and trajectory deviation change rate as fuzzy control input variables, and joint torque increment as the output variable. A fuzzy control rule base is built based on the robot joint's motion characteristics and historical control experience. The planned motion trajectory of each joint is input into the fuzzy controller, and parsed through three steps: fuzzification, fuzzy inference, and declarative analysis, to obtain the target angle, angular velocity, and torque of each joint in each control cycle.
[0035] This step achieves precise mapping from global task to joint motion through motion path decomposition. Combined with fuzzy control analysis, it eliminates the need to establish a complex and precise dynamic model for multi-degree-of-freedom robots, effectively solving the problems of low accuracy and poor adaptability of traditional PID control caused by strong coupling between joints and nonlinear characteristics. At the same time, it ensures the stability and safety of robot motion.
[0036] Step S400: Introduce a reinforcement learning mechanism to optimize the control strategy of the target multi-degree-of-freedom robot based on the M joint control parameters, determine the M joint control commands, and control the robot's operation through the M joint control commands.
[0037] Specifically, joint control commands are the final control signals generated after reinforcement learning optimization and sent directly to the joint servo driver. They are usually current commands or torque commands and are the final execution signals that drive the robot's movement.
[0038] Specifically, a Deep Deterministic Policy Gradient (DDPG) reinforcement learning agent is initialized. The robot's joint position state information, joint control parameters, trajectory tracking error, and environmental disturbances are used as the state space, while the activation weights and torque compensation amounts of the fuzzy control rules are used as the action space. A composite reward function is designed with the objectives of minimizing trajectory tracking error, maximizing motion stability, and minimizing energy consumption. M joint control parameters are input into the reinforcement learning agent as basic control quantities. The agent outputs corresponding optimization adjustment amounts based on the current state to dynamically compensate the output torque of the fuzzy controller, generating the final M joint control commands and sending them to the servo driver. After the robot executes the control commands, the actual running effect is fed back to the reinforcement learning agent. The agent calculates the reward value of this control based on the reward function, updates the neural network parameters, and continuously optimizes the control strategy.
[0039] This step, by introducing a reinforcement learning mechanism, enables online autonomous optimization of the control strategy, which can effectively compensate for unmodeled dynamic errors, load changes, and environmental interference, further improving the robot's trajectory tracking accuracy and significantly enhancing the system's adaptability to different working conditions and tasks.
[0040] Furthermore, joint state estimation is performed based on the M joint operation data and the robot global image to obtain M joint position state information, including: establishing M joint coordinate systems, a robot base coordinate system, and a global coordinate system; filtering and aligning the M joint operation data to obtain M usable joint operation data; using the M joint coordinate systems to estimate the joint state of the M usable joint operation data to obtain M initial joint position information; using the global coordinate system to perform joint detection and state estimation on the robot global image to obtain M joint global position information; and fusing and estimating the M initial joint position information and the M joint global position information based on the robot base coordinate system to obtain M joint position state information.
[0041] Specifically, the joint coordinate system refers to the local Cartesian coordinate system attached to each joint link. It serves as the reference for describing the motion of a single joint, with its origin located at the joint's rotation center and its Z-axis coinciding with the joint's rotation axis. The robot base coordinate system refers to the internal unified reference coordinate system fixed to the robot's base. It does not change with the robot's movement and serves as the final unified reference system for the position and orientation of all joints and end effectors. The global coordinate system refers to the absolute reference coordinate system fixed within the robot's workspace. It serves as the measurement reference for the vision sensors and provides an absolute position reference unaffected by the robot's own motion. Initial joint position information is joint position and orientation data independently estimated by the joint end sensor group, representing relative position information in the joint's local coordinate system. Global joint position information is joint position and orientation data independently detected by the global vision sensor, representing absolute position information in the global coordinate system with no accumulated error. Fusion estimation refers to the process of unifying the two independently estimated position information to the same reference through coordinate transformation, and then performing weighted calculations based on the real-time confidence levels of each sensor to obtain the optimal estimation result.
[0042] Specifically, a three-level coordinate system is established according to the DH parameter method, a standard in robotics. Each joint has its own joint coordinate system, with the origin of each system located at the rotation center of the corresponding joint, and the Z-axis coinciding with the joint's rotation axis. Simultaneously, the origin of the global coordinate system is set at the lower left corner of the workspace, with the Z-axis pointing vertically upwards and parallel to the base coordinate system. The synchronously acquired joint operation data is preprocessed, including using a first-order low-pass filter with a cutoff frequency of 50Hz to remove high-frequency vibration noise from the motor and reducer, and using an extended Kalman filter to remove random noise and abnormal jump values caused by electromagnetic interference. The filtered joint data is then aligned using a microsecond-level timestamp from a unified hardware synchronization clock. To address the issue that the 1kHz sampling rate of the joint sensor is much higher than the 30Hz sampling rate of the vision sensor, cubic spline interpolation is used to complete the joint data in the visual sampling gap, resulting in M usable joint operation data that are strictly time-aligned. The usable data is then processed using the joint coordinate system of each joint. The acceleration and angular velocity data are converted into first joint position information through Runge-Kutta numerical integration. This information is then compared with the rotational position data output by the motor encoder for error verification. If the comparison result is within the preset ±0.05° error threshold, the two data are fused through complementary filtering to obtain M initial joint position information. If the result exceeds the threshold, the encoder data is deemed invalid, and only the inertial integration result is used as the initial joint position information.
[0043] The robot's global image acquired synchronously is processed using a pre-established global coordinate system. The ArUco marker template matching algorithm is used to identify the feature markers of all joints in the image, and the pixel coordinates of each marker are extracted. Combined with the pre-calibrated intrinsic parameters of the vision sensor, including focal length, principal point coordinates, distortion coefficients and extrinsic parameters, the position and attitude of the camera in the global coordinate system are used to solve the PNP three-dimensional pose, generating M second joint position information. Then, through the transformation relationship between the global coordinate system and the camera coordinate system, it is converted into M joint global position information in the global coordinate system.
[0044] Based on the robot's base coordinate system, a fusion estimation is performed. The transformation matrix from the coordinate system of each joint to the base coordinate system is obtained by calculating the robot's forward kinematics. The transformation matrix from the global coordinate system to the base coordinate system is obtained through pre-calibration. These two sets of transformation matrices are used to uniformly transform the initial joint position information and the global joint position information to the base coordinate system. The fusion weight of the two data sources is dynamically adjusted according to the robot's current motion speed. The precise position state information of M joints in the base coordinate system is obtained by weighted averaging.
[0045] This step, by establishing a standardized three-level coordinate system, fundamentally solves the spatial alignment problem of multi-source heterogeneous sensor data, laying a unified mathematical foundation for subsequent fusion estimation. The adoption of a dual-redundant architecture with independent estimation at the joint end and the vision end enables mutual backup of sensors. When one sensor fails or its data becomes invalid, the other can still provide basic state feedback, significantly improving the system's robustness and reliability. The dynamic weighted fusion mechanism based on the base coordinate system fully leverages the complementary advantages of the joint sensor's high sampling rate and high dynamic response, and the vision sensor's lack of cumulative error and high absolute position accuracy. This effectively solves the problems of large cumulative error and inability to detect joint flexibility deformation and external impact deviations inherent in traditional single encoder control, providing a solid and reliable state feedback foundation for subsequent high-precision motion planning and closed-loop control.
[0046] Furthermore, the joint state is estimated using the M joint coordinate systems on the M available joint operating data to obtain M initial joint position information, including: obtaining M joint acceleration and angular velocity data and M motor rotation position data based on the M available joint operating data; performing numerical integration and joint position estimation on the M joint acceleration and angular velocity data using the M joint coordinate systems to obtain M first joint position information; comparing and analyzing the M first joint position information with the M motor rotation position data; if the comparison result is within a preset error threshold, performing joint state estimation by combining the M motor rotation position data and the M first joint position information to obtain M initial joint position information; if the comparison result exceeds the preset error threshold, using the M first joint position information as the M initial joint position information.
[0047] Specifically, motor rotational position data refers to the absolute / relative rotational angle data of the motor shaft output by the encoder installed on the non-load end of the motor, which is the theoretical rotational position of the joint after conversion by the reduction ratio of the reducer. Numerical integration is the process of recursively calculating the kinematic data output by the IMU, that is, integrating the angular velocity to obtain the joint angle increment, and integrating the acceleration to obtain the joint velocity and displacement increment, which is the core method of IMU attitude calculation. First joint position information refers to the joint position and attitude data obtained only through numerical integration of IMU data. It is characterized by high sampling rate and high dynamic response, but there is drift error that accumulates over time.
[0048] Specifically, from the obtained M available joint operation data, the three-dimensional acceleration data, three-dimensional angular velocity data, and motor rotation position data corresponding to each joint are extracted according to sensor type. For example, for the pitch joint of the boom of a 6-DOF automotive parts assembly industrial robot arm, the x, y, and z-axis accelerations and angular velocities at the joint's output end are obtained, along with the 17-bit absolute rotation angle data output by the encoder of the servo motor driving the joint. In the joint's own local coordinate system, the extracted angular velocity data is recursively integrated using the fourth-order Runge-Kutta numerical integration method to obtain the joint angle increment at each sampling time. After accumulation, the joint rotation angle is obtained. Simultaneously, the acceleration data is integrated to obtain the joint angle. The linear velocity and linear displacement of the joint are combined to obtain the first joint position information of the joint. For example, when the IMU sampling rate is 1kHz, the angular velocity in each sampling period is 0.5rad / s, and the integral yields an angle increment of 0.0005rad. Accumulating this angle value to the previous moment gives the joint angle at the current moment. The motor rotation position data is divided by the reduction ratio of the joint reducer to convert it into the theoretical rotation position of the joint. The difference between this theoretical position and the corresponding angle value in the first joint position information is calculated to obtain the real-time deviation. This deviation is then compared with a preset error threshold. If the comparison result is within the preset error threshold (e.g., the theoretical angle of the boom pitch joint is 30.02°), the real-time deviation obtained by IMU integration is considered acceptable. If the actual angle is 30.04° and the deviation is 0.02° (less than the ±0.05° threshold), a first-order complementary filtering algorithm is used to fuse the two. The encoder data corrects the low-frequency drift of the IMU, while the IMU data captures the high-frequency dynamic changes of the joint, obtaining the initial joint position information. If the comparison result exceeds a preset error threshold—for example, if the encoder data suddenly jumps to 35° due to gear slippage in the reducer, resulting in a deviation of 4.98°—the encoder data is deemed invalid, unreliable encoder data is automatically filtered out, and the first joint position information is directly used as the initial joint position information. The above calculation process is executed independently and in parallel in each joint's coordinate system, without interference, ensuring the system's real-time performance. In practical applications, when robots on an automotive welding production line are performing high-speed spot welding, the reducer gears may slip momentarily due to impact loads, causing abnormal jumps in encoder data. The system will automatically switch to pure IMU mode to continue the welding task and prevent workpiece scrapping. When a minimally invasive medical robot in an operating room is subjected to electromagnetic interference from an electrosurgical unit, the encoder data generates a large amount of noise, and the error exceeds the threshold. The system will automatically shield the encoder data and use only the IMU for state estimation to ensure surgical safety. When a collaborative robot has a minor collision with a human, the joints undergo flexible deformation. The encoder cannot detect this deformation, but the IMU can directly measure the actual joint position change, thereby adjusting the control strategy in a timely manner to avoid harm to the human body.
[0049] Numerical integration is a mathematical method that converts the instantaneous kinematic quantities output by the IMU into cumulative kinematic quantities that change over time. It is the core technology for independent state estimation by the joint-end IMU. It directly determines the accuracy of the first joint position information and the dynamic response speed. Its calculation error will directly affect the error verification results of subsequent data with encoder data and the final fusion estimation accuracy.
[0050] The basic principle of angular velocity integration is that the angular velocity sensor (gyroscope) outputs the instantaneous angular velocity (unit: rad / s) of the joint around the three axes of its own coordinate system. By integrating over time, the angle increment of the joint on the corresponding axis can be obtained, and after accumulation, the absolute rotation angle of the joint can be obtained.
[0051] Continuous domain formula: ; represents the angle of the joint at any given moment, which is equal to the initial angle plus the sum of the time increments of all instantaneous angular velocities from the initial moment to the current moment. Wherein, It is the absolute rotation angle of the joint at time t, in radians or degrees, and is the joint position state quantity that we ultimately need to solve for; It is the absolute rotation angle of the joint at the initial moment (t=0), and the unit is radians or degrees; It is a definite integral operation, which means that the physical quantity in parentheses is calculated by time accumulation from the initial time 0 to the current time t. yes The instantaneous angular velocity of the joint at any given moment, measured in radians per second, is the gyroscope's angular velocity at any given time. The output rotational speed of the joint about its own rotation axis is a function that changes continuously with time; It is the integration variable (time), with the unit being seconds, representing all intermediate moments traversed during the integration process, from 0 to t; It is the time differential, with the unit of seconds, representing an infinitesimally small time interval, and is the basic unit of continuous integration.
[0052] Discrete domain formula: This represents the joint angle at the current sampling moment, which is equal to the angle at the previous sampling moment, plus the angular velocity at the current moment multiplied by the sampling period. Here, k is the sampling moment number, an integer starting from 0 and incrementing. k=0 corresponds to the initial moment, k=1 corresponds to the first sampling moment, and so on. It is the absolute rotation angle of the joint at the kth sampling time, in radians or degrees, and is calculated every 1ms. That is, k=100 corresponds to the joint angle at t=100ms. It is the absolute rotation angle of the joint at the (k-1)th sampling time (the previous time), in radians or degrees, and is the result obtained from the previous calculation; It is the instantaneous angular velocity of the joint at the k-th sampling time, in radians per second; it is the gyroscope measurement value output by the IMU at the k-th sampling time, updated every 1ms. It is the sampling period, in seconds, which is the time interval between two adjacent sampling moments. The IMU sampling rate is 1kHz, so Δt = 0.001s.
[0053] The fourth-order Runge-Kutta method, or RK4, calculates the slope at four distinct points within the integration interval, takes a weighted average as the average slope, and then multiplies it by the sampling period to obtain the angle increment. The calculation formula is as follows: ; ; ; ; ; Among them, It is the instantaneous angular velocity at the start of the sampling period, which is directly taken from the IMU measurement value at the previous moment, corresponding to the first approximate slope of the Euler method; It is the predicted slope at the midpoint of the sampling period, using Predict the angular velocity after half a cycle to initially correct the error in the slope at the starting point; It is the corrected slope at the midpoint of the sampling period, using The midpoint angular velocity was recalculated to further improve the accuracy of the midpoint slope; It is the predicted slope at the end of the sampling period, using Predict the angular velocity after the entire cycle and capture the motion changes at the end of the cycle. The coefficients 1, 2, 2, 1 are the optimal weights derived from Taylor expansion, and the midpoint slope... and The weights are higher because they are more representative of the average motion state over the entire period. The equivalent average angular velocity over the entire sampling period is obtained by weighting the four slopes, and then multiplying it by the sampling period. The angle increment is obtained and accumulated to the angle at the previous moment. Get the angle at the current moment For linearly changing angular velocities, there is almost no truncation error.
[0054] The fourth-order Runge-Kutta method is far more accurate than the Euler and trapezoidal methods, while the computational cost is only about twice that of the trapezoidal method, and it can run in real time at a sampling frequency of 1kHz. It can accurately capture the instantaneous angular velocity and acceleration changes during high-speed joint movements, making it suitable for scenarios involving rapid start-stop and high-speed movement of robots.
[0055] The basic principle of acceleration integration is that the accelerometer outputs the instantaneous linear acceleration (unit: m / s²) of the joint around the three axes of its own coordinate system. The linear velocity of the joint is obtained by the first integration, and the linear displacement of the joint is obtained by the second integration.
[0056] Linear velocity integral formula: ; represents the linear velocity of the joint at any given moment, which is equal to the initial velocity plus the sum of the time increments of all instantaneous accelerations from the initial moment to the current moment. Wherein It is the linear velocity of the joint at time t, in meters per second, and is the instantaneous velocity of the joint output end at time t. It is the initial linear velocity of the joint, measured in meters per second, when the robot is initially stationary. =0; yes The actual motion acceleration of the joint at any given time, measured in meters per second², is the net motion acceleration after subtracting the gravitational acceleration component from the accelerometer reading.
[0057] Linear velocity-displacement integral formula: ; represents the linear displacement of the joint at any given moment, which is equal to the initial displacement plus the time accumulation of the initial velocity, plus the quadratic time accumulation of all instantaneous accelerations. yes The linear displacement of the joint at any given time, measured in meters, is the spatial displacement of the joint's output end relative to its initial position. It is the linear displacement of the joint at the initial moment, in meters. The position of the robot is usually set to 0 when it is initialized. It is a double definite integral operation, which means first integrating the acceleration to obtain the velocity, and then integrating the velocity to obtain the displacement; It is the inner integral variable (time), with the unit being seconds. It is the time variable of the inner integral in a double integral, ranging from 0 to... .
[0058] Linear velocity integral formula: This indicates that the current velocity is equal to the velocity at the previous moment plus the current acceleration multiplied by the sampling period. It is the first The linear velocity of the joint at each sampling time point, in meters per second, and the instantaneous velocity of the joint is calculated every 1 ms; It is the linear velocity of the joint at the (k-1)th sampling time, in meters per second, and is the velocity value obtained in the previous calculation; It is the actual motion acceleration of the joint at the kth sampling time, in meters per second², and is the acceleration sensor measurement value after gravity compensation.
[0059] Linear displacement integral formula: This indicates that the displacement at the current moment is equal to the displacement at the previous moment plus the current velocity multiplied by the sampling period. It is the linear displacement of the joint at the kth sampling time, in meters, and the instantaneous displacement of the joint is calculated every 1ms; It is the linear displacement of the joint at the (k-1)th sampling time, in meters, and is the displacement value obtained in the previous calculation.
[0060] The continuous time t corresponds to the sampling time index k in the discrete domain, k=0, 1, 2, ... corresponding to t=0, 1ms, 2ms, ... respectively; the infinitesimal time interval The sampling period Δt corresponding to the discrete domain is 0.001s; continuous function , Discrete sampling sequence corresponding to the discrete domain , This refers to the IMU output sequence with one data point every 1ms; the definite integral operation corresponds to the discrete domain accumulation operation, that is, the incremental accumulation of each sampling period.
[0061] This step fundamentally solves the problem that traditional single-encoder control cannot detect the deviation between actual and theoretical positions caused by joint flexibility deformation, transmission backlash, gear wear, and external impacts by constructing a dual-redundant state estimation architecture at the joint end, with direct measurement by IMU and indirect measurement by encoder. Through a real-time error verification mechanism between encoder and IMU data, automatic detection and isolation of encoder faults are achieved, significantly improving the fault tolerance and reliability of joint state estimation and avoiding robot loss of control due to the failure of a single sensor. The complementary filtering fusion algorithm fully combines the advantages of high dynamic response of IMU and high short-term accuracy of encoder, effectively suppressing the cumulative drift error of pure IMU integration. The architecture of independent parallel computing for each joint greatly reduces the computational complexity of the system, making it suitable for real-time operation on embedded controllers and providing a reliable local state foundation for subsequent high-precision fusion with global vision data.
[0062] Furthermore, the global coordinate system is used to perform joint detection and state estimation on the global image of the robot to obtain M joint global position information, including: pre-establishing M joint feature templates; performing joint matching detection on the global image of the robot according to the M joint feature templates to obtain M joint pose features; combining the calibration parameters of the vision sensor to perform state estimation on the M joint pose features to generate M second joint position information; and using the global coordinate system to perform coordinate transformation on the M second joint position information to obtain M joint global position information.
[0063] Specifically, joint feature templates refer to pre-collected and standardized digital templates containing unique visual identifiers for each joint. These are typically artificially generated tags with unique IDs, such as ArUco tags, or the geometric contour features of the joint itself, serving as a benchmark for rapid and accurate joint identification. Visual sensor calibration parameters are divided into intrinsic and extrinsic parameters. Intrinsic parameters include camera focal length, principal point coordinates, and lens distortion coefficients, describing the camera's optical characteristics. Extrinsic parameters include the camera's position and orientation in the global coordinate system, describing the spatial relationship between the camera and the global coordinate system. Second joint position information refers to the three-dimensional position and orientation data of the joints in the camera coordinate system obtained through visual image processing; it is an intermediate result of independent state estimation at the vision end. Global joint position information refers to the absolute position and orientation data of the joints obtained by converting the second joint position information to the global coordinate system. This data is unaffected by the robot's own motion and has no accumulated error.
[0064] Specifically, a feature template library for M joints is pre-established. A unique ArUco marker with a unique ID is affixed to the non-occluded surface of each joint. Marker images are acquired under different angles and illumination intensities. After standardization, a template library containing marker features is generated, and the precise position of each marker in the corresponding joint coordinate system is recorded. The synchronously acquired global robot images are pre-processed, including grayscale conversion, Gaussian filtering for noise reduction, and histogram equalization to enhance contrast, resulting in clear pre-processed images. Then, the ArUco marker detection algorithm is used to perform frame-by-frame matching and detection on the pre-processed images according to the pre-established feature templates, identifying the ArUco markers of all joints. The pixel coordinates and rotation angles of the four corner points of each marker are extracted to obtain the pose features of the M joints. Combining the calibrated intrinsic and extrinsic parameters of the vision sensor, the PNP algorithm is used to perform 3D pose calculation on the pose features of each joint, generating the position information of the M second joints in the camera coordinate system. Using the pre-calibrated camera extrinsic parameter matrix, the position information of the second joints in the camera coordinate system is transformed to the global coordinate system, obtaining the global position information of the M joints.
[0065] This step achieves rapid and accurate joint identification through a matching detection method based on pre-established feature templates; combined with the PNP 3D pose calculation method based on visual sensor calibration parameters, it can accurately recover the 3D position and orientation of the joint from the 2D image; the global position information of the joint obtained by transforming to the global coordinate system provides an absolute position reference, fundamentally eliminating the cumulative error problem of the joint end sensor, and providing reliable global correction data for subsequent multi-source fusion state estimation.
[0066] Furthermore, based on the robot base coordinate system, the position information of the M initial joints and the global position information of the M joints are fused and estimated to obtain the position state information of the M joints. This includes: obtaining the coordinate transformation relationship matrices between the M joint coordinate systems and the global coordinate system and the robot base coordinate system, respectively; using the coordinate transformation relationship matrices to perform coordinate transformation on the position information of the M initial joints and the global position information of the M joints to obtain the transformed joint position information and the transformed global position information; and performing weighted fusion estimation on the transformed joint position information and the transformed global position information to obtain the position state information of the M joints.
[0067] Specifically, transforming joint position information refers to mapping the initial joint position information in the joint coordinate system to position and pose data in the robot's base coordinate system through a corresponding coordinate transformation matrix, thus preserving the high dynamic response characteristics of the joint sensors. Transforming global position information refers to mapping the global joint position information in the global coordinate system to position and pose data in the robot's base coordinate system through a corresponding coordinate transformation matrix, thus preserving the absolute position characteristics of the vision sensors, which have no accumulated errors.
[0068] Specifically, two sets of coordinate transformation relationship matrices are obtained. The transformation matrix from the M joint coordinate systems to the robot base coordinate system is obtained by forward kinematics calculation using the robot standard DH parameter method. The 4×4 homogeneous transformation matrix from the joint coordinate system to the base coordinate system is calculated based on the pre-calibrated DH parameters. The transformation matrix from the global coordinate system to the robot base coordinate system is obtained through one-time pre-calibration. Specifically, a 12×9 checkerboard calibration board is placed simultaneously in the common field of view of the two coordinate systems, and the pose of the calibration board in the two coordinate systems is measured respectively. The transformation relationship between the two is obtained by solving the least squares method.
[0069] These two sets of transformation matrices are used to perform coordinate transformation on M initial joint position information and M global joint position information, respectively. Homogeneous matrix multiplication is used to unify the two types of position and pose data into the robot's base coordinate system, resulting in M transformed joint position information and M transformed global position information. The fusion weights of the two types of data are dynamically adjusted according to the robot's current motion speed, sensor data quality, and environmental conditions. For example, when the robot is static or moving at low speed, the visual sensor data is stable and has no accumulated error, so the weight is set to 0.7, and the joint sensor weight is set to 0.3. When the robot is moving at high speed, the high sampling rate of 1kHz of the joint sensor has obvious advantages, so the weight is set to 0.8, and the visual sensor weight is set to 0.2. When occlusion or blurring of the visual image is detected, causing the confidence level to fall below the preset threshold of 0.6, the visual weight is automatically reduced to 0, and only the transformed joint position information is used. When encoder data failure is detected, causing low confidence of the initial joint position information, the joint sensor weight is automatically reduced to 0, and only the transformed global position information is used. Finally, the final position state information of the M joints in the base coordinate system is calculated by weighted averaging.
[0070] Preferably, all coordinate transformations are implemented using a 4×4 homogeneous transformation matrix, unifying rotations and translations in three-dimensional space into a single mathematical matrix. The standard form of the matrix is: ; The 3×3 submatrix R in the upper left corner is the rotation matrix, describing the pose difference between the two coordinate systems; the 3×1 vector in the upper right corner... : is a translation vector, describing the positional difference between the origins of two coordinate systems; the last row is fixed as [0, 0, 0, 1], used to implement homogeneous coordinate operations.
[0071] The transformation matrix from the joint coordinate system to the robot's base coordinate system, also known as the DH parameter method, is a standard method in robotics for calculating coordinate transformations between links. It uses four DH parameters to describe the spatial relationship between adjacent links, and finally obtains the transformation matrix from any joint coordinate system to the base coordinate system through matrix multiplication. The DH parameters are defined as follows: For the i-th joint and the i-th link, the four DH parameters are as follows: Linkage length :ith The vertical distance from the axis of one joint to the axis of the i-th joint.
[0072] Linkage torsion angle :ith One joint axis around Rotate to the angle of the i-th joint axis.
[0073] Joint offset :ith The distance from the origin of the link coordinate system to the origin of the i-th link coordinate system along the axis of the i-th joint.
[0074] Joint angle :ith The angle by which a link coordinate system rotates about the i-th joint axis to the i-th link coordinate system.
[0075] The formula for the transformation matrix of adjacent links, i-th Transformation matrix from one joint coordinate system to the i-th joint coordinate system i is:
[0076] The transformation matrix from the global coordinate system to the robot's base coordinate system is calculated using the calibration method. This transformation matrix describes the spatial relationship between the global coordinate system fixed in the workspace and the base coordinate system fixed on the robot's base. It is obtained through a single calibration. The calibration principle is to use a common calibration object as an intermediate reference, measure its pose in both coordinate systems, and then solve for the transformation matrix. Specific calculation steps: Prepare the calibration object: Use a high-precision 12×9 checkerboard calibration board, with each square having a side length of 20mm. Obtain the pose of the calibration board in the global coordinate system. The calibration board is placed in the robot's workspace, and images of the calibration board are captured using a global vision sensor. Combining the camera's intrinsic and extrinsic parameters, the PNP algorithm is used to calculate the pose matrix of the calibration board in the global coordinate system. Obtain the pose of the calibration board in the robot's base coordinate system. The robot's end effector is controlled to sequentially contact at least four non-collinear corner points on the calibration plate, and the coordinates of each contact point in the robot's base coordinate system are recorded. Then, the pose matrix of the calibration plate in the base coordinate system is fitted using the least squares method. Solve for the transformation matrix from the global coordinate system to the base coordinate system. According to the chain rule of coordinate transformation, ,therefore ,in, for The inverse matrix.
[0077] This step fundamentally solves the problem of inconsistent data space between different sensors by establishing a unified robot base coordinate system as the fusion benchmark for multi-source data, achieving seamless alignment between local data at the joint end and global data at the vision end. A dynamic weighted fusion mechanism is adopted, fully leveraging the complementary advantages of the joint sensors' high sampling rate and high dynamic response, and the vision sensors' lack of cumulative error and high absolute accuracy. Simultaneously, the weights can be adaptively adjusted according to real-time operating conditions and sensor states, significantly improving the accuracy and robustness of the fusion estimation. Through an automatic weight adjustment mechanism under abnormal conditions, automatic detection and fault-tolerant handling of sensor failures are achieved. When one sensor fails, the other can still provide reliable status feedback, preventing the robot from going out of control.
[0078] Furthermore, by combining the robot's working environment map and task requirements, motion path decomposition and fuzzy control analysis are performed on the position and state information of the M joints to obtain M joint control parameters. This includes: performing global path planning for the target multi-degree-of-freedom robot based on the robot's working environment map and task requirements to determine the robot's global task path; obtaining the M joint operation constraints based on the mechanical characteristic information of the M joints; and performing motion path decomposition and fuzzy control analysis on the position and state information of the M joints based on the M joint operation constraints according to the robot's global task path to obtain the M joint control parameters.
[0079] Specifically, the robot's global task path refers to the continuous motion trajectory of the robot's end effector in Cartesian space, generated by global path planning. It contains a series of ordered path points, each corresponding to the end effector's three-dimensional position and orientation. Mechanical characteristic information refers to the inherent physical parameters of each joint of the robot, including maximum angular velocity, maximum angular acceleration, maximum output torque, range of motion, reducer reduction ratio, and moment of inertia, determined by the robot's mechanical design. Operational constraints are limitations derived from the joint's mechanical characteristics and task safety requirements, which must be met during joint movement. They are prerequisites for generating a legal motion trajectory and prevent joint movement exceeding limits, which could lead to equipment damage or safety accidents.
[0080] Specifically, it loads a pre-built 3D raster working environment map constructed via SLAM and the input specific task requirements; it then combines the environment map and task requirements to employ an improved A / B algorithm. The algorithm performs global path planning, generating a collision-free global task path for the robot's end effector from its current position to the screw feeder gripping point and then to the battery module mounting hole. The path contains 100 uniformly distributed Cartesian space path points, each containing the end effector's 3D coordinates and quaternion pose. It also reads the mechanical characteristics of M joints pre-stored in the robot controller, such as the maximum angular velocity of the boom pitch joint being 30° / s, the maximum angular acceleration being 50° / s², and the maximum output torque being 50N. The range of motion is -30° to +120°. The operating constraints of each joint are obtained through parameter conversion, including velocity constraints, acceleration constraints, torque constraints and position constraints.
[0081] Starting with the current position and state information of M joints, the motion path is decomposed according to the global task path, under the premise of satisfying all joint operation constraints. A numerical inverse kinematics algorithm is used to convert each Cartesian space path point into a corresponding joint space angle sequence, while avoiding robot singularities, generating a smooth running trajectory for M joints. The running trajectory and current actual position of each joint are input into a fuzzy controller for analysis. The joint trajectory deviation and trajectory deviation change rate are selected as input variables for fuzzy control, and the joint torque increment is selected as the output variable. A Mamdani-type fuzzy control rule base is constructed based on the motion characteristics of the robot joints and historical control experience. For example, "if the trajectory deviation is positive and the deviation change rate is positive, then the output torque increment is positive". Through three steps of fuzzification, fuzzy inference, and decanting using the center of gravity method, the control parameters such as the target angle, angular velocity, and output torque of each joint in each 1ms control cycle are obtained.
[0082] The fuzzy controller adopted is a Mamdani-type fuzzy controller, whose core structure consists of four parts: a fuzzification interface, a fuzzy rule base, a fuzzy inference engine, and a declarative interface. The fuzzification interface is responsible for converting precise input quantities into fuzzy linguistic variables. The fuzzy rule base stores control rules based on expert experience and historical data. The fuzzy inference engine performs logical reasoning based on the fuzzy input quantities and the rule base to obtain fuzzy output quantities. The declarative interface converts the fuzzy output quantities into executable precise control parameters. During construction, the input and output variables and the universe of discourse are first determined. Joint trajectory deviation *e* and trajectory deviation change rate *ec* are selected as input variables, and the universe of discourse is set to [-3, 3]. The joint torque increment is selected. As the output variable, the universe of discourse is set to [-2,2]. Next, fuzzy subsets are defined, dividing both input and output variables into seven fuzzy subsets: {negative large, negative medium, negative small, zero, positive small, positive medium, positive large}. Triangular membership functions are used to describe the membership degree of each fuzzy subset, with narrow triangles used for the zero subset to improve control accuracy and trapezoids for the two extreme subsets to enhance robustness. Then, a fuzzy rule base is constructed, based on the robot's joint motion characteristics and a large amount of control experimental data, generating the rule "If e is A and ec is B, then...". The standard control rules are in the form of "C", such as "if the deviation is positive and the rate of change of the deviation is positive, then the output torque increment is positive". Finally, the Mamdani minimum-maximum inference method is selected for fuzzy inference, and the center of gravity method is used for defuzzification processing to convert the fuzzy output into an accurate torque increment value. This controller does not require the establishment of an accurate dynamic model of the robot joint and can effectively adapt to the strong coupling and nonlinear characteristics between the joints of multi-degree-of-freedom robots.
[0083] This step, through a hierarchical and progressive architecture of global path planning, joint constraint extraction, motion decomposition, and fuzzy analysis, first ensures the global collision-free nature and task feasibility of robot motion. Second, it avoids equipment damage caused by exceeding motion limits through strict joint operation constraints. Finally, through fuzzy control analysis that does not require a precise dynamic model, it effectively solves the problems of low accuracy and poor adaptability of traditional PID control caused by strong coupling and nonlinear characteristics between joints in multi-degree-of-freedom robots. This ensures the smoothness and safety of motion and can adapt to the needs of various application scenarios such as industrial assembly, medical surgery, and human-robot collaboration.
[0084] Furthermore, based on the robot's global task path and the running constraints of the M joints, motion path decomposition and fuzzy control analysis are performed on the position and state information of the M joints to obtain M joint control parameters. This includes: decomposing the motion path of the M joints based on the robot's global task path and the running constraints of the M joints to obtain the planned running trajectory of the M joints; selecting joint fuzzy control input variables and output variables, and constructing a fuzzy control rule base according to the joint motion characteristics of the target multi-degree-of-freedom robot; and performing fuzzy control analysis on the planned running trajectory of the M joints based on the fuzzy control rule base, the joint fuzzy control input variables, and the output variables to obtain the M joint control parameters.
[0085] Specifically, motion path decomposition refers to the process of converting the global task path of the robot's end effector in Cartesian space into an independent time-series motion trajectory for each joint in joint space, while satisfying the physical limits of each joint, using an inverse kinematics algorithm. It serves as a bridge connecting global planning and joint control. The fuzzy control input variables are the input signals of the fuzzy controller, consisting of the trajectory deviation *e* between the actual and planned joint positions, and the rate of change *ec* of the trajectory deviation, which together reflect the real-time tracking state of the joint. The fuzzy control output variables are the output signals of the fuzzy controller, representing the joint torque increment. Precise control of joint movement is achieved by adjusting the output torque of the servo motor.
[0086] Specifically, starting with the current position and state information of M joints and ending with the robot's global task path, the motion path is decomposed under the premise of satisfying the obtained running constraints of the M joints. A numerical inverse kinematics algorithm is used to convert each Cartesian space path point on the global task path into a corresponding joint space angle sequence. Simultaneously, the robot's singularities are avoided using the Jacobian matrix pseudo-inverse method. Then, a fifth-order polynomial interpolation algorithm is used to smooth the discrete angle sequences, generating a continuous, impact-free running trajectory for each joint. Next, a fuzzy controller adapted to the target robot is constructed, selecting the joint trajectory deviation e and the trajectory deviation change rate ec as input variables, setting their universe of discourse to [-3,3]. The joint torque increment is selected as the input variable. As output variables, their universe of discourse is set to [-2,2]. All three variables are divided into seven fuzzy subsets: {negative large, negative medium, negative small, zero, positive small, positive medium, positive large}. A triangular membership function is used to describe the membership degree of each fuzzy subset. The zero subset uses a narrow triangle to improve control accuracy under small deviations, while the two-ended subsets use a trapezoid to enhance robustness under large deviations. Then, based on the robot's joint motion characteristics, a standard Mamdani-type fuzzy control rule base is constructed. Finally, the planned trajectory of each joint and the real-time acquired actual position are input into the fuzzy controller for analysis. First, the precise input is converted into fuzzy linguistic variables using the membership function. Then, the min-max inference method is used to logically infer the fuzzy output based on the fuzzy rule base. Finally, the centroid method is used to convert the fuzzy output into precise joint torque increment values. Combined with the joint's basic torque, complete control parameters such as the target angle, angular velocity, and output torque for each joint within each 1ms control cycle are obtained.
[0087] This step, through motion path decomposition under constraints, fundamentally avoids equipment damage and safety accidents caused by joint movement exceeding limits, and the generated smooth trajectory ensures the stability of robot motion; through a customized fuzzy controller, there is no need to build a complex and accurate dynamic model of a multi-degree-of-freedom robot, effectively solving the problems of low accuracy and poor adaptability of traditional PID control caused by strong coupling between joints and nonlinear characteristics.
[0088] Furthermore, a reinforcement learning mechanism is introduced to optimize the control strategy of the target multi-degree-of-freedom robot based on the M joint control parameters, and to determine the M joint control commands. This includes: controlling the target multi-degree-of-freedom robot to perform task execution monitoring based on the M joint control parameters to obtain the current robot feedback state; and introducing a reinforcement learning mechanism to optimize the control strategy based on the current robot feedback state to determine the M joint control commands.
[0089] Specifically, the current robot feedback status refers to the multi-dimensional data collected in real time by the sensor group and vision sensor during the robot's task execution, which reflects the robot's actual operating status and task completion status. This includes the actual joint position, speed, torque, trajectory tracking error, force on the end effector, and environmental interference information.
[0090] Specifically, the target multi-degree-of-freedom robot is driven by M joint control parameters. Simultaneously, a task execution monitoring module is activated to collect real-time robot operation data through installed sensor groups and vision sensors. This includes the actual position, velocity, output torque, and tracking error of each joint relative to the planned trajectory, as well as the force conditions of the end effector and environmental interference information. This data is then integrated to obtain the current robot feedback state. Next, a pre-simulated and pre-trained DDPG reinforcement learning agent is initialized. The current robot feedback state is used as the agent's state space input, and the activation weight adjustment and joint torque compensation of the fuzzy control rules are used as the action space output. A composite reward function is designed with the objectives of minimizing trajectory tracking error, maximizing motion stability, and minimizing energy consumption. The trajectory tracking error weight accounts for 60%, motion stability weight for 30%, and energy consumption weight for 10%. A positive reward is given when the trajectory tracking error is less than 0.05 mm, and a penalty is given when the error exceeds 0.2 mm or the joint vibration amplitude exceeds a threshold.
[0091] The agent dynamically compensates the basic joint control parameters based on the current feedback state and the output of the pre-trained policy network, generating the final M joint control commands and sending them to the servo driver for execution. After the robot executes the control commands, it feeds back the new operating state to the agent. The agent calculates the reward value of this action based on the reward function, updates the parameters of the critic network and the actor network through the experience playback mechanism, and performs a policy update every 100 control cycles to continuously optimize the control strategy.
[0092] This step, by introducing the DDPG reinforcement learning mechanism of offline pre-training and online fine-tuning, realizes the autonomous online optimization of the control strategy. It can effectively compensate for problems that are difficult to solve by traditional control methods, such as unmodeled dynamic errors of the robot, load changes, frictional drift and environmental interference. It enhances the system's adaptability to different working conditions and tasks and extends the robot's service life.
[0093] Furthermore, a reinforcement learning mechanism is introduced to optimize the control strategy for the current robot feedback state, determining M joint control commands. This includes: defining the robot state space, joint motion space, and reward function based on the reinforcement learning mechanism and the multi-objective of robot joint control; constructing a control strategy reinforcement algorithm model based on the robot state space, joint motion space, and reward function; and using the control strategy reinforcement algorithm model to optimize the control strategy for the current robot feedback state, determining M joint control commands.
[0094] Specifically, the robot state space refers to the set of all robot operating states that the reinforcement learning agent can perceive. It serves as the basis for the agent's decisions and includes multi-dimensional continuous variables such as joint position, velocity, output torque, trajectory tracking error, end effector force, joint vibration amplitude, and environmental disturbances. The joint motion space refers to the set of all control adjustment actions that the reinforcement learning agent can output. It consists of the activation weight adjustment and joint torque compensation amounts of the fuzzy control rules, both of which are continuous values and do not require changes to the core structure of the original fuzzy control. The control policy reinforcement algorithm model is an end-to-end control optimization model built on Deep Deterministic Policy Gradient (DDPG), including an actor network and a critic network, capable of learning the optimal control policy in the continuous motion space. The experience replay mechanism refers to storing the states, actions, rewards, and next state data generated by the agent's interaction with the environment in an experience pool, randomly sampling samples for model training, breaking data correlation, and improving training stability.
[0095] Specifically, based on the reinforcement learning mechanism and the multi-objective requirements of robot joint control, the core elements of the model are defined. The state space is defined as a multi-dimensional continuous vector, containing the position error, velocity error, output torque, and three-dimensional force on the end effector for each joint. The action space is defined as a 50-dimensional continuous vector, containing the activation weight adjustments for 49 fuzzy control rules and one global torque compensation, with an adjustment range of [-0.2, 0.2] to ensure control stability. The reward function is defined as a composite weighted reward function, with the following formula: ; in A reward is given for trajectory tracking accuracy; the smaller the error, the higher the reward. A penalty of -10 is given when the error exceeds 0.2mm. For motion stability rewards, the smaller the change in joint acceleration, the higher the reward. For energy consumption rewards, the smaller the output torque, the higher the reward. A control strategy reinforcement algorithm model is constructed based on the DDPG algorithm. The model contains four deep neural networks: an actor network, a critic network, a target actor network, and a target critic network. Each network contains three hidden layers with 64 neurons per layer. An experience replay pool with a capacity of 100,000 is also constructed to store interaction data. Then, a robot simulation platform is used to generate 10,000 sets of control data under different loads and speeds for offline pre-training, enabling the model to initially grasp the basic control laws. Finally, the pre-trained model is deployed to the robot controller. The current robot feedback state is input into the model, and the actor network outputs the corresponding action adjustment amount to dynamically compensate the basic joint control parameters, generating the final M joint control commands and sending them to the servo driver for execution. After the robot executes the commands, it generates a new state and reward, which are stored in the experience replay pool. Every 100 control cycles, 64 samples are randomly selected to fine-tune the model online, update the network parameters, and continuously optimize the control strategy.
[0096] For the control strategy enhancement algorithm model, an offline pre-training mode combining simulation environment-driven and experience playback training is adopted. Basic training of the DDPG control strategy model is completed before the actual deployment of the robot, enabling the model to initially grasp the basic laws of joint control. This avoids the equipment damage and safety risks that may occur during online training and significantly shortens the time for subsequent online fine-tuning. Pre-training first constructs a digital twin simulation environment based on the target robot's 3D CAD model and measured dynamic parameters, accurately reproducing joint rotational inertia, friction coefficient, reducer clearance, and sensor noise. Simultaneously, multiple sets of diverse training conditions covering different loads, speeds, trajectories, and random disturbances are generated. The DDPG model, including actors, commentators, and corresponding target networks, is initialized, and hyperparameters such as the learning rate of 1e-4 and the discount factor of 0.99 are set. During training, a task is randomly selected to initialize the simulation state. The actor network outputs fuzzy rule weight adjustments and torque compensation, interacting with the environment to generate a "state-action-reward-next state" quadruple, which is stored in an experience pool with a capacity of 100,000. Once the data volume reaches the target, 64 samples are randomly selected in each batch. The critic network is updated using mean squared error, the actor network is updated using policy gradient, and the target network is softly updated. Iterative training continues until the average reward is stable for 10 consecutive rounds and the trajectory tracking error is ≤ ±0.1mm, at which point the model converges. Finally, 2000 sets of untrained working conditions are used to validate the model.
[0097] This step constructs a DDPG reinforcement learning model adapted to robot joint control by defining a multi-objective composite reward function and a continuous action space. The training method, combining offline pre-training with online fine-tuning, ensures both the initial control performance of the model and continuous autonomous optimization of the control strategy. This model effectively compensates for problems that traditional control methods struggle to address, such as unmodeled dynamic errors, load variations, frictional drift, and environmental disturbances. This further improves the robot's trajectory tracking accuracy and enhances the system's adaptability to different working conditions and tasks, allowing it to adapt to various application scenarios without manual parameter readjustment.
[0098] In summary, the joint control method for a multi-degree-of-freedom robot provided in this application has the following technical effects: 1. By using multi-sensor data acquisition and hierarchical state estimation, path planning, intelligent control and optimization, the shortcomings of traditional robot control, such as reliance on a single sensor, easy error accumulation and weak dynamic adaptability, are effectively overcome. The overall positioning accuracy and operational stability of joint movements are improved, enabling the robot to reliably complete its tasks under complex working conditions.
[0099] 2. By combining local joint estimation with global visual estimation, multiple independent solutions for joint positions are achieved, which are then fused together in a unified base coordinate system, significantly improving the reliability and accuracy of position estimation results. Multiple data streams mutually verify and complement each other, effectively suppressing single-sensor drift and providing stable and accurate state input for subsequent high-precision control.
[0100] 3. By combining the environmental map with task requirements, a global collision-free path is planned, and trajectory decomposition is performed under joint mechanical constraints to ensure robot motion safety and within physical limits. Fuzzy control analysis is employed to reduce reliance on precise dynamic models and improve control adaptability to nonlinear, strongly coupled joint systems.
[0101] Example 2, based on the same inventive concept as the joint control method for multi-degree-of-freedom robots in the foregoing examples, such as... Figure 2 As shown in the figure, this application provides a joint control system for a multi-degree-of-freedom robot, the system comprising: The sensor assembly installation module 11 is used to install sensor assemblies on M joints of the target multi-degree-of-freedom robot. The sensor assemblies integrate accelerometers, angular velocity sensors, and motor encoders. A vision sensor is also installed within the workspace of the target multi-degree-of-freedom robot. The joint position information acquisition module 12 is used to collect the running data of the M joints through the sensor assemblies and simultaneously acquire the global image of the robot collected by the vision sensor. Based on the running data of the M joints and the global image of the robot, joint state estimation is performed to obtain the position state information of the M joints. The joint control parameter acquisition module 13 is used to perform motion path decomposition and fuzzy control analysis on the position state information of the M joints in conjunction with the robot's working environment map and task requirements to obtain the control parameters of the M joints. The joint control command determination module 14 is used to introduce a reinforcement learning mechanism to optimize the control strategy of the target multi-degree-of-freedom robot based on the control parameters of the M joints, determine the control commands of the M joints, and control the robot's operation through the control commands of the M joints.
[0102] Furthermore, the joint position information acquisition module 12 is also used to perform the following steps: establishing M joint coordinate systems, a robot base coordinate system, and a global coordinate system; filtering and aligning the M joint operation data to obtain M usable joint operation data; using the M joint coordinate systems to estimate the joint state of the M usable joint operation data to obtain M initial joint position information; using the global coordinate system to perform joint detection and state estimation on the robot global image to obtain M global joint position information; and fusing and estimating the M initial joint position information and the M global joint position information based on the robot base coordinate system to obtain M joint position state information.
[0103] Furthermore, the joint position information acquisition module 12 is also used to perform the following steps: based on the M available joint operation data, obtain M joint acceleration and angular velocity data and M motor rotation position data; use the M joint coordinate systems to perform numerical integration and joint position estimation on the M joint acceleration and angular velocity data to obtain M first joint position information; compare and analyze the M first joint position information with the M motor rotation position data; if the comparison result is within a preset error threshold, combine the M motor rotation position data and the M first joint position information to perform joint state estimation to obtain M initial joint position information; if the comparison result exceeds the preset error threshold, use the M first joint position information as the M initial joint position information.
[0104] Furthermore, the joint position information acquisition module 12 is also used to perform the following steps: pre-establish M joint feature templates, perform joint matching detection on the robot global image according to the M joint feature templates to obtain M joint posture features; combine the calibration parameters of the vision sensor to perform state estimation on the M joint posture features to generate M second joint position information; use the global coordinate system to perform coordinate transformation on the M second joint position information to obtain M joint global position information.
[0105] Furthermore, the joint position information acquisition module 12 is also used to perform the following steps: acquiring the coordinate transformation relationship matrices between the M joint coordinate systems and the global coordinate system and the robot base coordinate system, respectively; using the coordinate transformation relationship matrices to perform coordinate transformation on the M initial joint position information and the M joint global position information to obtain M transformed joint position information and M transformed global position information; and performing weighted fusion estimation on the M transformed joint position information and M transformed global position information to obtain M joint position state information.
[0106] Furthermore, the joint control parameter acquisition module 13 is also used to perform the following steps: perform global path planning for the target multi-degree-of-freedom robot by combining the robot's working environment map and task requirements, and determine the robot's global task path; obtain the M joint operation constraints based on the mechanical characteristic information of the M joints; and perform motion path decomposition and fuzzy control analysis on the position state information of the M joints according to the robot's global task path and the M joint operation constraints to obtain the M joint control parameters.
[0107] Furthermore, the joint control parameter acquisition module 13 is also used to perform the following steps: decompose the motion path of the M joint position state information according to the global task path of the robot based on the running constraints of the M joints, and obtain the running planning trajectory of the M joints; select the joint fuzzy control input variables and output variables, and construct a fuzzy control rule base according to the joint motion characteristics of the target multi-degree-of-freedom robot; perform fuzzy control parsing on the running planning trajectory of the M joints based on the fuzzy control rule base, the joint fuzzy control input variables and output variables, and obtain the M joint control parameters.
[0108] Furthermore, the joint control command determination module 14 is also used to perform the following steps: control the target multi-degree-of-freedom robot to perform task execution monitoring based on the M joint control parameters, and obtain the current robot feedback state; introduce a reinforcement learning mechanism to optimize the control strategy for the current robot feedback state, and determine the M joint control commands.
[0109] Furthermore, the joint control command determination module 14 is also used to perform the following steps: defining the robot state space, joint motion space, and reward function according to the reinforcement learning mechanism and the multi-objective of robot joint control; constructing a control strategy reinforcement algorithm model based on the robot state space, joint motion space, and reward function; and using the control strategy reinforcement algorithm model to optimize the control strategy of the current robot feedback state to determine M joint control commands.
[0110] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A joint control method for a multi-degree-of-freedom robot, characterized in that, The method includes: Sensor groups are installed on M joints of the target multi-degree-of-freedom robot. The sensor groups integrate accelerometers, angular velocity sensors and motor encoders. At the same time, vision sensors are set in the workspace of the target multi-degree-of-freedom robot. The sensor group collects the running data of M joints, and at the same time acquires the global image of the robot collected by the vision sensor. Based on the running data of the M joints and the global image of the robot, the joint state is estimated to obtain the position state information of the M joints. By combining the robot's working environment map and task requirements, motion path decomposition and fuzzy control analysis are performed on the position and state information of the M joints to obtain the control parameters of the M joints. A reinforcement learning mechanism is introduced to optimize the control strategy of the target multi-degree-of-freedom robot based on the M joint control parameters, determine the M joint control commands, and control the robot's operation through the M joint control commands.
2. The joint control method for a multi-degree-of-freedom robot as described in claim 1, characterized in that, Based on the running data of the M joints and the global image of the robot, joint state estimation is performed to obtain the position state information of the M joints, including: Establish M joint coordinate systems, a robot base coordinate system, and a global coordinate system; The running data of the M joints are filtered and aligned to obtain the running data of the M available joints. The joint state of the M available joints is estimated using the coordinate system of the M joints to obtain the position information of the M initial joints. Using the global coordinate system, joint detection and state estimation are performed on the global image of the robot to obtain global position information of M joints; Based on the robot's base coordinate system, the initial joint position information and the global joint position information of the M joints are fused and estimated to obtain the position state information of the M joints.
3. The joint control method for a multi-degree-of-freedom robot as described in claim 2, characterized in that, Using the M joint coordinate systems, the joint state is estimated from the operational data of the M available joints to obtain M initial joint position information, including: Based on the operating data of the M available joints, M joint acceleration and angular velocity data and M motor rotation position data are obtained; Numerical integration and joint position estimation are performed on the acceleration and angular velocity data of the M joints using the M joint coordinate systems to obtain the position information of the M first joints. The M first joint position information is compared and analyzed with the M motor rotation position data. If the comparison result is within the preset error threshold, the joint state is estimated by combining the M motor rotation position data and the M first joint position information to obtain M initial joint position information. If the comparison result exceeds the preset error threshold, the M first joint position information will be used as the M initial joint position information.
4. The joint control method for a multi-degree-of-freedom robot as described in claim 2, characterized in that, Using the global coordinate system, joint detection and state estimation are performed on the global image of the robot to obtain global position information for M joints, including: M joint feature templates are pre-established, and joint matching detection is performed on the global image of the robot according to the M joint feature templates to obtain M joint posture features; By combining the calibration parameters of the vision sensor, the state of the M joint pose features is estimated to generate M second joint position information; The global coordinate system is used to perform coordinate transformation on the position information of the M second joints to obtain the global position information of the M joints.
5. The joint control method for a multi-degree-of-freedom robot as described in claim 2, characterized in that, Based on the robot's base coordinate system, the initial joint position information and the global joint position information of the M joints are fused and estimated to obtain the position state information of the M joints, including: Obtain the coordinate transformation relationship matrices between the M joint coordinate systems and the global coordinate system and the robot base coordinate system, respectively; The coordinate transformation relationship matrix is used to transform the M initial joint position information and the M joint global position information to obtain M transformed joint position information and M transformed global position information; The M joint position information and the M global position information are weighted and fused to obtain the M joint position state information.
6. The joint control method for a multi-degree-of-freedom robot as described in claim 1, characterized in that, By combining the robot's working environment map and task requirements, the position and state information of the M joints is decomposed into motion paths and analyzed using fuzzy control techniques to obtain the control parameters for the M joints, including: By combining the robot's working environment map and task requirements, global path planning is performed on the target multi-degree-of-freedom robot to determine the robot's global task path; Based on the mechanical characteristics of the M joints, the operating constraints of the M joints are obtained; Based on the robot's global task path and the running constraints of the M joints, the motion path is decomposed and fuzzy control is analyzed to obtain the control parameters of the M joints.
7. The joint control method for a multi-degree-of-freedom robot as described in claim 6, characterized in that, Based on the robot's global task path and the running constraints of the M joints, motion path decomposition and fuzzy control analysis are performed on the position and state information of the M joints to obtain the control parameters of the M joints, including: Based on the robot's global task path and the running constraints of the M joints, the motion path is decomposed into the position and state information of the M joints to obtain the running planning trajectory of the M joints. Select the joint fuzzy control input and output variables, and construct a fuzzy control rule base based on the joint motion characteristics of the target multi-degree-of-freedom robot; Based on the fuzzy control rule base, the joint fuzzy control input variables and output variables, the running trajectories of the M joints are analyzed by fuzzy control to obtain the control parameters of the M joints.
8. The joint control method for a multi-degree-of-freedom robot as described in claim 1, characterized in that, A reinforcement learning mechanism is introduced to optimize the control strategy of the target multi-degree-of-freedom robot based on the M joint control parameters, and to determine the M joint control commands, including: Based on the M joint control parameters, the target multi-degree-of-freedom robot is controlled to perform task execution monitoring and obtain the current robot feedback status. A reinforcement learning mechanism is introduced to optimize the control strategy based on the current robot feedback state, and to determine M joint control commands.
9. The joint control method for a multi-degree-of-freedom robot as described in claim 8, characterized in that, A reinforcement learning mechanism is introduced to optimize the control strategy based on the current robot feedback state, determining M joint control commands, including: Based on reinforcement learning mechanisms and multi-objective robot joint control, the robot state space, joint motion space, and reward function are defined. Based on the robot's state space, joint motion space, and reward function, a control strategy reinforcement algorithm model is constructed. The control strategy is enhanced by the aforementioned control strategy to optimize the control strategy of the current robot feedback state, and to determine M joint control commands.
10. A joint control system for a multi-degree-of-freedom robot, characterized in that, The system is used to perform a joint control method for a multi-degree-of-freedom robot according to any one of claims 1 to 9, the system comprising: A sensor group mounting module is used to mount sensor groups on M joints of a target multi-degree-of-freedom robot. The sensor groups integrate accelerometers, angular velocity sensors and motor encoders, and a vision sensor is also set in the workspace of the target multi-degree-of-freedom robot. The joint position information acquisition module is used to collect the running data of M joints through the sensor group, and at the same time acquire the global image of the robot collected by the vision sensor. Based on the running data of the M joints and the global image of the robot, the module performs joint state estimation to obtain the position state information of the M joints. The joint control parameter acquisition module is used to perform motion path decomposition and fuzzy control analysis on the position and state information of the M joints by combining the robot's working environment map and task requirements, and to obtain the M joint control parameters. The joint control command determination module is used to introduce a reinforcement learning mechanism to optimize the control strategy of the target multi-degree-of-freedom robot based on the M joint control parameters, determine the M joint control commands, and control the robot's operation through the M joint control commands.