High-altitude operation six-degree-of-freedom mechanical arm and self-adaptive anti-interference control method and system thereof

Through the carbon fiber topology optimization design and diversion channel structure, combined with the dynamic disturbance compensation and impedance control of DDPG and TD3 algorithms, wind resistance and load adjustment are optimized, and the problems of traditional robotic arms being disturbed by wind and load changes in high-altitude operations are solved, real-time obstacle avoidance path planning with high precision and low energy consumption are achieved.

CN120287301AActive Publication Date: 2025-07-11BEIJING FEIYU TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510610665.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-11
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The operation robotic arms of traditional flying robots are susceptible to strong wind interference during high altitude operations, resulting in dislocation of the end position and vibrations when the load changes. The path planning method does not consider joint energy consumption, and the obstacle avoidance algorithm cannot deal with dynamic obstacles, which affects operating accuracy, stability and efficiency.

Method used

The arm body and diversion channel structure designed with carbon fiber topology optimization is adopted, combined with the dynamic perturbation observer of the DDPG algorithm and the impedance control of the TD3 algorithm, and real-time obstacle avoidance paths are generated based on the multi-objective PPO path planning algorithm, and wind resistance and load adjustment are optimized to deal with dynamic obstacles.

Benefits of technology

It significantly improves the wind resistance, stability and load adaptability of the robotic arm, reduces energy consumption, improves operating accuracy and obstacle avoidance efficiency, and meets the needs of narrow scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120287301A_ABST
    Figure CN120287301A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle control, in particular to an aerial work six-degree-of-freedom mechanical arm and a self-adaptive anti-interference control method and system.The self-adaptive anti-interference control method comprises the steps that a mechanical arm structure is manufactured through an arm body designed in a carbon fiber topological optimization mode and a joint shell provided with a flow guide groove; constructing a dynamic disturbance observer based on a DDPG algorithm, outputting a dynamic compensation amount through the dynamic disturbance observer, calculating a wind power disturbance value in combination with a disturbance compensation algorithm, and integrating the wind power disturbance value to a control law of the unmanned aerial vehicle; adjusting impedance parameters on line based on a TD3 algorithm; and generating a real-time obstacle avoidance path based on an optimized multi-target PPO path planning algorithm, and executing a path planning task of the corresponding real-time obstacle avoidance path. The unmanned aerial vehicle has the advantages that the wind disturbance resistance is remarkably enhanced, the load sudden change stability is remarkably optimized, the task energy consumption is remarkably reduced, the narrow space adaptability is remarkably improved, and the light weight is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle control, and particularly to a six-degree-of-freedom robotic arm for high-altitude operations, its adaptive disturbance rejection control method and system. Background Art

[0002] With the gradual in-depth application of unmanned aerial vehicles in the fields of power equipment inspection, wind turbine blade maintenance, overhead line maintenance, etc., the working robotic arms of traditional flying robots still have obvious limitations in many aspects. For example, in the high-altitude wind turbine blade maintenance operation scenario, the influence of strong wind on the working robotic arm is a major challenge. The traditional working robotic arm is vulnerable to strong wind interference, resulting in the deviation of the end pose, which in turn affects the operation accuracy, working quality and efficiency of the working robotic arm. This is mainly because the structural design of the traditional working robotic arm is insufficient in dealing with strong wind. There is no reasonable flow guiding design for its joint housing, resulting in a large wind resistance. At the same time, the arm body is relatively heavy, and the inertial impact generated by the wind load easily changes the end pose of the working robotic arm. Another example is that during the overhead line maintenance process, when the working robotic arm carries tools of different weights at high altitude, the load will suddenly change. Due to the lack of an effective load change recognition and response mechanism, the traditional working robotic arm is prone to large vibrations, which will not only affect the operation stability, but may also cause damage to the working robotic arm itself. At the same time, in some scenarios of long-term continuous operation, since the traditional path planning method only focuses on the path length and ignores the joint energy consumption, the traditional working robotic arm needs to perform joint movements frequently. If the joint energy consumption is not considered, it will lead to waste of energy and increase the usage cost. In addition, in some complex working environments, most of the existing obstacle avoidance algorithms rely on offline calculations and are difficult to cope with dynamic obstacles. For example, in the bridge maintenance site, there may be the appearance of dynamic obstacles such as personnel and vehicles. The traditional obstacle avoidance algorithms cannot react in time, easily resulting in collisions between the working robotic arm and obstacles, affecting work safety and efficiency.

[0003] Therefore, the present application provides a six-degree-of-freedom robotic arm for high-altitude operations, its adaptive disturbance rejection control method and system. Summary of the Invention

[0004] Based on this, it is necessary to provide a six-degree-of-freedom robotic arm for high-altitude operations, its adaptive disturbance rejection control method and system for the above technical problems.

[0005] According to the first aspect of the present invention, an adaptive disturbance rejection control method for a six-degree-of-freedom manipulator for high-altitude operations is provided, including: manufacturing a manipulator structure with an arm body designed by carbon fiber topology optimization and a joint housing provided with a flow guiding groove; constructing a dynamic disturbance observer based on the DDPG algorithm, outputting a dynamic compensation amount through the dynamic disturbance observer, combining with a disturbance compensation algorithm, calculating a wind disturbance value and integrating it into the control law of the unmanned aerial vehicle for real-time dynamic disturbance rejection compensation; online adjusting impedance parameters based on the TD3 algorithm to cope with sudden load changes; generating a real-time obstacle avoidance path based on an optimized multi-objective PPO path planning algorithm and performing a path planning task for the corresponding real-time obstacle avoidance path to cope with dynamic obstacles.

[0006] Optionally, the arm body designed by carbon fiber topology optimization is an internal truss structure generated by multi-condition topology optimization of the arm body based on finite element analysis software, and the internal truss structure is made of carbon fiber composite material. Among them, the load conditions for multi-condition topology optimization include a wind pressure of 15 m / s and an end load of 10 kg, the constraint conditions are that the maximum stress ≤ 300 MPa and the first-order modal frequency ≥ 50 Hz, and the objective function includes minimizing the mass of the arm body and maximizing the stiffness of the arm body.

[0007] Optionally, the flow guiding groove is arranged on the surface of the joint housing and has a V-shaped structure. Geometric parameters including groove depth, inclination angle, and spacing are defined, and the wind resistance coefficient under a wind pressure condition of 15 m / s is simulated and verified by CFD software to optimize the aerodynamic performance.

[0008] Optionally, constructing a dynamic disturbance observer based on the DDPG algorithm, outputting a dynamic compensation amount through the dynamic disturbance observer, combining with a disturbance compensation algorithm, calculating a wind disturbance value and integrating it into the control law of the unmanned aerial vehicle, includes: collecting real-time wind pressure data, joint state data, and historical error data to construct a state space data set; constructing and training the DDPG algorithm, and constructing a dynamic disturbance observer based on the trained DDPG algorithm; where the DDPG algorithm consists of a first Actor network for outputting a dynamic compensation amount and a first Critic network for evaluating the value of the dynamic compensation amount. Both the first Actor network and the first Critic network are composed of a three-layer fully connected structure, and the output nodes of each layer of the fully connected structure are 256, 128, and 64 respectively; during the training process of the DDPG algorithm, inputting the state space data set, outputting the dynamic compensation amount in the corresponding state space, calculating the loss value in the real-time state space according to the mean square tracking error loss function, and online updating the DDPG algorithm parameters in combination with the real-time wind pressure data to obtain an online updated DDPG algorithm; based on the online updated DDPG algorithm, outputting the dynamic compensation amount in the real-time state space, combining with a disturbance compensation algorithm, calculating a wind disturbance value and integrating it into the control law of the unmanned aerial vehicle for real-time dynamic disturbance rejection compensation.

[0009] Optionally, the real-time wind pressure data is collected by a plurality of pressure sensors disposed on the surface of the robotic arm structure. The joint state data includes the joint angle, joint speed, and current loop torque feedback by the joint encoder. The historical error data is the end pose tracking error data within the defined historical control period.

[0010] Optionally, the formula of the disturbance compensation algorithm is: ; where is the wind disturbance value, is the dynamic compensation amount, is the state space data set, and are obtained after calibration by the particle swarm optimization algorithm. Among them, takes the value of 12, takes the value of 0.5, is the control error.

[0011] Optionally, the online adjustment of the impedance parameters based on the TD3 algorithm includes: collecting the real-time end contact force, joint torque deviation data, and trajectory error data, and constructing an observation state data set; among them, the end contact force includes three force components and three torque components collected by a six-axis force sensor. The three force components are Fx, Fy, and Fz respectively, and the three torque components are Mx, My, and Mz respectively. The joint torque deviation data is the difference between the actual torque of each joint and the expected value. The trajectory error data includes the end position deviation value and the end speed deviation value; constructing and training the TD3 algorithm; where the TD3 algorithm consists of a second Actor network for outputting stiffness and damping and a second Critic network for preventing overestimation. The second Actor network is composed of a four-layer fully connected structure, and the output nodes of each layer of the fully connected structure are 512, 256, 128, and 64 respectively. The second Critic network adopts a double Q network structure; when training the TD3 algorithm, input the observation state data set and combine the defined OU noise to obtain the trained TD3 algorithm; based on the trained TD3 algorithm, output the stiffness and damping under the real-time observation state and use them as the impedance parameters under the real-time observation state; based on the impedance parameters under the real-time observation state, perform PID control and adjust the corresponding impedance parameters to cope with sudden changes in the load.

[0012] Optionally, the optimized multi-objective PPO path planning algorithm generates a real-time obstacle avoidance path and executes the path planning task of the corresponding real-time obstacle avoidance path, including: collecting real-time environmental perception data, where the real-time environmental perception data includes real-time point cloud data collected by a binocular vision camera and IMU attitude data collected by an IMU attitude sensor, performing downsampling processing on the real-time point cloud data to obtain a depth map of 128×128 pixels, and inputting it into a three-layer convolutional neural network for feature extraction to output dynamic obstacle features, and performing state encoding on the dynamic obstacle features and IMU attitude data to obtain a state encoding result; optimizing the multi-objective PPO path planning algorithm through a defined dynamic reward function and a policy update mechanism, where the formula of the dynamic reward function is: ; where R is the total reward, , , are weight coefficients used to balance the importance of each reward term, is the path completion degree, indicating the completion degree of the current path, with a value range between 0 and 1, is the energy consumption, indicating the energy consumed by the robotic arm structure in the operation task, and the calculation method is the sum of the squares of the joint torques, and the specific formula is: , is the torque of the th joint, is the total number of joints, is the collision risk, indicating the reciprocal of the distance between the robotic arm structure and the nearest obstacle and performing normalization processing; constructing and training a policy network structure, where the policy network structure consists of an LSTM layer and a fully connected layer for outputting the probability distribution of the motion of each joint. During the training process of the policy network structure, input the state encoding result, output the probability distribution of the motion of each joint in the corresponding environmental state, and calculate the total reward in the corresponding environmental state in combination with the dynamic reward function, and update the parameters of the policy network structure based on the total reward in the corresponding environmental state to obtain an optimized multi-objective PPO path planning algorithm; real-time detecting dynamic obstacles in the environment through a binocular vision camera, and based on the corresponding state encoding result, triggering a local path replanning mechanism based on a defined local path update time to replan the local path, based on the optimized multi-objective PPO path planning algorithm, outputting the probability distribution of the motion of each joint under the corresponding local path, and selecting the optimal joint motion, and taking the optimal joint motion as the real-time obstacle avoidance path; in response to the real-time obstacle avoidance path, executing the path planning task of the corresponding real-time obstacle avoidance path.

[0013] According to the second aspect of the present invention, an adaptive disturbance rejection control system for a six-degree-of-freedom manipulator for high-altitude operations is provided, including: a structure manufacturing module for manufacturing the manipulator structure using a carbon fiber topologically optimized designed arm body and a joint housing provided with a flow guiding groove; a real-time dynamic disturbance rejection compensation module for constructing a dynamic disturbance observer based on the DDPG algorithm, outputting a dynamic compensation amount through the dynamic disturbance observer, combining with a disturbance compensation algorithm, calculating the wind disturbance value and integrating it into the control law of the unmanned aerial vehicle for real-time dynamic disturbance rejection compensation; an impedance parameter adjustment module for online adjusting the impedance parameters based on the TD3 algorithm to cope with sudden load changes; an obstacle avoidance path planning module for generating a real-time obstacle avoidance path based on an optimized multi-objective PPO path planning algorithm and executing the path planning task of the corresponding real-time obstacle avoidance path to cope with dynamic obstacles.

[0014] According to the third aspect of the present invention, a six-degree-of-freedom manipulator for high-altitude operations is provided, including a manipulator structure manufactured using the above-mentioned adaptive disturbance rejection control method.

[0015] The advantages and beneficial effects of the present invention are as follows: 1. The wind disturbance resistance ability is significantly enhanced: Through the structural design and optimization of the flow guiding groove, the wind resistance coefficient can be effectively reduced. Combining with the lightweight characteristics of the carbon fiber arm body, the inertial impact of the wind load on the manipulator structure can be greatly reduced, and the end positioning accuracy can be improved. At the same time, combined with the dynamic disturbance observer, by real-time estimating the wind disturbance value and performing feedforward compensation, the end positioning error under the wind pressure condition of 15 m / s can be effectively reduced, and the compensation response delay time can be effectively reduced; 2. The stability of sudden load changes is significantly optimized: Using a six-axis force sensor to fuse the joint encoder data to achieve millisecond-level synchronous feedback of the end contact force and joint torque, thereby improving the sensitivity of identifying load changes. At the same time, combined with the dynamic impedance control technology, the impedance parameters can be adjusted in real time according to the load mass, and thus the stabilization time and vibration amplitude when the load is suddenly applied can be effectively reduced; 3. The energy consumption efficiency is improved breakthroughly: By optimizing the multi-objective PPO path planning algorithm and introducing an energy consumption reward term and an energy consumption weight coefficient into the dynamic reward function, the task energy consumption can be significantly reduced, thereby improving the energy utilization efficiency. At the same time, combined with the arm body designed by carbon fiber topology optimization, by reducing the arm body inertia, the ineffective energy consumption of the joint motor during the acceleration stage can be further reduced; 4. The adaptability to narrow spaces is significantly improved: Through the synergistic effect of online real-time obstacle avoidance and compact joint design, not only can a new path be quickly generated and responded to, but also the joint axial dimension of the manipulator structure can be reduced, so that the minimum clearance that the manipulator structure can pass through is ≤200 mm, meeting the requirements of narrow scenarios such as bridge maintenance. Description of the Drawings

[0016] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the following detailed description of the preferred embodiments. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Also, throughout the drawings, the same reference symbols are used to represent the same components. In the drawings:

[0017] Figure 1 It is a flowchart of an embodiment of an adaptive disturbance rejection control method for a six-degree-of-freedom robotic arm for high-altitude operations according to the present application.

[0018] Figure 2 It is a comparative curve graph of an embodiment of a load mutation PID control response according to the present application.

[0019] Figure 3 It is a comparative curve graph of an embodiment of the energy consumption of flight operations according to the present application.

[0020] Figure 4 It is a schematic structural diagram of an embodiment of an adaptive disturbance rejection control system for a six-degree-of-freedom robotic arm for high-altitude operations according to the present application.

[0021] Figure 5 It is a schematic structural diagram of an embodiment of a robotic arm structure according to the present application.

[0022] Figure 6 It is a schematic structural diagram of an embodiment of an arm body according to the present application.

[0023] Wherein, 100, robotic arm structure; 200, diversion groove; 300, joint; 400, joint housing; 500, arm body. Detailed Embodiments

[0024] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0025] Embodiment 1

[0026] Refer to the attached Figure 1 , the present application provides an adaptive disturbance rejection control method for a six-degree-of-freedom robotic arm for high-altitude operations, and the method includes the following steps.

[0027] S1: Fabricate the robotic arm structure using an arm body with carbon fiber topology optimization design and a joint housing with a flow guiding groove.

[0028] Specifically, most traditional robotic arms use aluminum alloy or steel arm bodies, and their defects are as follows: (1) High inertia moment: Traditional materials have a high density (steel: 7.8 g / cm³), restricting the dynamic response speed; (2) High drag coefficient: Without optimized aerodynamic design, the drag coefficient ≥ 0.8, and the additional power consumption increases by 30% under strong winds. Therefore, this application fabricates the robotic arm structure using an arm body with carbon fiber topology optimization design and a joint housing with a flow guiding groove. Among them, an internal truss structure is generated through finite element topology optimization, reducing its mass by 40% and increasing its stiffness by 25%. Combining with the flow guiding groove design of the joint housing and defining geometric parameters including groove depth, inclination angle, and spacing, through CFD software simulation and verification, the drag coefficient under a wind pressure condition of 15 m / s is reduced from 0.8 to 0.35, and the joint drive power consumption under the same operation task is reduced by 28%.

[0029] Furthermore, the arm body with carbon fiber topology optimization design is an internal truss structure generated after multi-condition topology optimization of the arm body based on finite element analysis software, and the internal truss structure is made of carbon fiber composite material. Among them, the load conditions for multi-condition topology optimization include a wind pressure of 15 m / s and a terminal load of 10 kg, the constraint conditions are maximum stress ≤ 300 MPa and first-order modal frequency ≥ 50 Hz, and the objective function includes minimizing the mass of the arm body and maximizing the stiffness of the arm body.

[0030] Furthermore, the flow guiding groove is arranged on the surface of the joint housing and has a V-shaped structure. Define geometric parameters including groove depth, inclination angle, and spacing, and verify the drag coefficient under a wind pressure condition of 15 m / s through CFD software simulation to optimize the aerodynamic performance; preferably, the groove depth is 3 mm, the inclination angle is 15°, and the spacing is 10 mm.

[0031] S2: Construct a dynamic disturbance observer based on the DDPG algorithm, output a dynamic compensation amount through the dynamic disturbance observer, combine with the disturbance compensation algorithm, calculate the wind disturbance value and integrate it into the control law of the unmanned aerial vehicle for real-time dynamic disturbance rejection compensation.

[0032] Specifically, when working at high altitudes, traditional robotic arms mostly adopt fixed feedforward compensation or an equivalent degree-of-freedom model based on a preset constant wind load (such as CN117325143A). Its defects are as follows: (1) Dependence on an offline wind field model: It is unable to sense dynamic wind speed changes (such as gusts, turbulence, etc.) in real time, resulting in compensation lag; (2) Sensitive to high-frequency disturbances: Fixed gain parameters are prone to oscillations under high-frequency random disturbances. Under the wind speed condition of 10 m / s, the measured end positioning error ≥ ±1.2 mm. Therefore, this application adopts a dynamic disturbance observer based on reinforcement learning, uses the DDPG algorithm to output real-time dynamic compensation amounts, combines with the disturbance compensation algorithm, and optimizes the wind disturbance value in real time to achieve feedforward compensation, which can dynamically adapt to wind speed changes, reduce the compensation delay to within 10 ms. At the same time, under the wind speed condition of 15 m / s, the end positioning error ≤ ±0.3 mm, and the disturbance rejection accuracy is improved by 60%.

[0033] Further, a dynamic disturbance observer is constructed based on the DDPG algorithm. The dynamic compensation amount is output through the dynamic disturbance observer, and combined with the disturbance compensation algorithm, the wind disturbance value is calculated and integrated into the control law of the unmanned aerial vehicle, including: collecting real-time wind pressure data, joint state data, and historical error data, and constructing a state space data set; constructing and training the DDPG algorithm, and constructing a dynamic disturbance observer based on the trained DDPG algorithm; among them, the DDPG algorithm consists of a first Actor network for outputting the dynamic compensation amount and a first Critic network for evaluating the value of the dynamic compensation amount. Both the first Actor network and the first Critic network are composed of a three-layer fully connected structure, and the output nodes of each layer of the fully connected structure are 256, 128, and 64 respectively; during the training process of the DDPG algorithm, the state space data set is input, the dynamic compensation amount in the corresponding state space is output, the loss value in the real-time state space is calculated according to the mean square tracking error loss function, and the DDPG algorithm parameters are updated online in combination with the real-time wind pressure data to obtain the online updated DDPG algorithm; based on the online updated DDPG algorithm, the dynamic compensation amount in the real-time state space is output, combined with the disturbance compensation algorithm, the wind disturbance value is calculated and integrated into the control law of the unmanned aerial vehicle for real-time dynamic disturbance rejection compensation.

[0034] Further, the real-time wind pressure data is collected by multiple pressure sensors arranged on the surface of the robotic arm structure. For example, the real-time wind pressure data can be collected by 16 pressure sensors (such as pressure sensors of model MPX5700AP) arranged on the surface of the robotic arm structure; the joint state data includes the joint angle, joint speed, and current loop torque fed back by the joint encoder; the historical error data is the end pose tracking error data within the defined historical control period. For example, the end pose tracking errors in the past 5 control periods.

[0035] Further, in the construction and training process of the DDPG algorithm, the present application constructs an initial DDPG algorithm with a first Actor network and a first Critic network. Among them, the first Actor network is used to output a dynamic compensation amount. , the first Critic network is used to evaluate the value of the dynamic compensation amount, specifically to evaluate the state-action value. The first Actor network and the first Critic network have the same structure and are both composed of three fully connected structures. The output nodes of each fully connected structure are 256, 128, and 64 respectively. The output results of each layer will be non-linearly transformed through the ReLU activation function to enhance the expression ability of the model. In the initial training stage of the DDPG algorithm, 10-hour wind pressure data, joint state data, and historical error data within the range of 5-20 m / s of wind speed can be collected to construct a state space data set. Subsequently, the constructed state space data set is input into the initial DDPG algorithm for training, and the loss value in the real-time state space is calculated according to the mean square tracking error loss function. The DDPG algorithm parameters are updated online in combination with the real-time wind pressure data (such as updating the DDPG algorithm parameters once every 1 hour), and an online updated DDPG algorithm is obtained.

[0036] Further, the formula of the perturbation compensation algorithm is: ; in the formula, is the wind force perturbation value, is the dynamic compensation amount, is the state space data set, and are obtained after being calibrated by the particle swarm optimization algorithm. Among them, takes the value of 12, takes the value of 0.5, is the control error.

[0037] Further, based on the online updated DDPG algorithm, the real-time state space data set is input, and the dynamic compensation amount in the real-time state space is output. In combination with the perturbation compensation algorithm, the wind force perturbation value is calculated and integrated into the control law of the unmanned aerial vehicle to achieve feedforward compensation for real-time dynamic disturbance compensation.

[0038] S3: Online adjust the impedance parameters based on the TD3 algorithm to cope with sudden load changes.

[0039] Specifically, in the prior art, an impedance control framework is mostly adopted to achieve force / position hybrid control by setting fixed stiffness parameters (such as CN110653805A). Specifically: in the task constraint space, an impedance model is defined based on fixed inertia, damping, and stiffness matrices. By calibrating the impedance parameter library offline under different loads and calling the preset parameters when switching loads, its defects are as follows: (1) Parameter curing leads to a lag in response when the load changes suddenly (settling time > 2s); (2) The change in environmental stiffness (such as the difference in softness and hardness of the contact surface) is not considered, which is likely to cause overshoot or under-compensation. Therefore, this application adopts an adaptive impedance control technology driven by the TD3 algorithm. By inputting a real-time observation state data set, the stiffness and damping are output in real time, and the stiffness and damping output in real time are used as the impedance parameters under the real-time observation state. Through PID control and execution of adjusting the corresponding impedance parameters, it is used to cope with sudden load changes. The settling time when a 5 kg-class load is suddenly applied can be shortened from 1.2 s of traditional PID control to 0.5 s (refer to Appendix Figure 2 ), the vibration amplitude decays by 80%, and the contact force tracking error ≤ ±3 N, so as to adapt to the sudden load change scenarios in different working environments.

[0040] Furthermore, the impedance parameters are adjusted online based on the TD3 algorithm, including: collecting real-time end contact force, joint torque deviation data, and trajectory error data to construct an observation state data set; among them, the end contact force includes three force components and three torque components collected by a six-dimensional force sensor. The three force components are Fx, Fy, and Fz respectively, and the three torque components are Mx, My, and Mz respectively. The joint torque deviation data is the difference between the actual torque of each joint and the expected value. The trajectory error data includes the end position deviation value and the end velocity deviation value; constructing and training the TD3 algorithm; among them, the TD3 algorithm consists of a second Actor network for outputting stiffness and damping and a second Critic network for preventing overestimation. The second Actor network is composed of a four-layer fully connected structure, and the output nodes of each layer of the fully connected structure are 512, 256, 128, and 64 respectively. The second Critic network adopts a double Q-network structure; when training the TD3 algorithm, input the observation state data set and combine the defined OU noise to obtain the trained TD3 algorithm; based on the trained TD3 algorithm, output the stiffness and damping under the real-time observation state and use them as the impedance parameters under the real-time observation state; based on the impedance parameters under the real-time observation state, through PID control and execution of adjusting the corresponding impedance parameters, it is used to cope with sudden load changes.

[0041] Furthermore, real-time end contact force, joint torque deviation data, and trajectory error data are collected. Among them, the end contact force can be collected by a six-axis force sensor (such as a six-axis force sensor of model ATI Gamma) for three force components and three torque components. The three force components are Fx, Fy, and Fz respectively, and the three torque components are Mx, My, and Mz respectively; the joint torque deviation data is the difference between the actual torque of each joint and the expected value . The trajectory error data includes the end position deviation value and the end velocity deviation value.

[0042] Furthermore, during the construction and training process of the TD3 algorithm, the present application constructs an initial TD3 algorithm using a second Actor network and a second Critic network. Among them, the second Actor network is used to output stiffness and damping . . The second Actor network is composed of a four-layer fully connected structure. The output nodes of each layer of the fully connected structure are 512, 256, 128, and 64 respectively. The second Critic network adopts a double Q network structure to prevent overestimation. At the same time, during the training process, exploration noise is introduced, specifically OU noise. Preferably, θ (the rate of mean regression) in the OU noise is set to 0.15, and σ (volatility, that is, the degree of perturbation) is set to 0.2 to accelerate the model convergence process.

[0043] Furthermore, based on the trained TD3 algorithm, the stiffness and damping in the real-time observation state are output and used as the impedance parameters in the real-time observation state (such as updating the impedance parameters every 10 ms); based on the impedance parameters in the real-time observation state, PID control is used to adjust the corresponding impedance parameters to cope with sudden changes in the load.

[0044] S4: Based on the optimized multi-objective PPO path planning algorithm, a real-time obstacle avoidance path is generated, and the path planning task of the corresponding real-time obstacle avoidance path is executed to cope with dynamic obstacles.

[0045] Specifically, in the prior art, obstacle avoidance algorithms based on geometric rules are mostly adopted. For example, the improved FM algorithm proposed in CN110653805A (specifically using a static cost function) is used to plan a collision-free path. Its defects are as follows: (1) Lack of dynamic obstacle avoidance ability: The planning depends on a static environment model and cannot respond to dynamic obstacles (such as drifting debris) in real time; (2) Energy consumption is not optimized: The path length priority strategy ignores the change of joint torque, resulting in a 30% increase in ineffective energy consumption. Therefore, the multi-objective PPO path planning algorithm involved in this application has been specifically adjusted in the design of the dynamic reward function and the policy update mechanism. By introducing an energy consumption reward term and an energy consumption weight coefficient into the dynamic reward function, the sum of the squares of the joint torques after optimization is reduced by 42%, and the overall energy consumption under the same operation task is reduced by 35% (refer to Appendix Figure 3 ) and can also shorten the response time and effectively improve the speed of obstacle avoidance path planning.

[0046] Furthermore, based on the optimized multi-objective PPO path planning algorithm, a real-time obstacle avoidance path is generated, and the path planning task of the corresponding real-time obstacle avoidance path is executed, including: collecting real-time environment perception data, where the real-time environment perception data includes real-time point cloud data collected by a binocular vision camera and IMU attitude data collected by an IMU attitude sensor, downsampling the real-time point cloud data to obtain a depth map of 128×128 pixels, and inputting it into a three-layer convolutional neural network for feature extraction to output dynamic obstacle features, and performing state encoding on the dynamic obstacle features and IMU attitude data to obtain a state encoding result; optimizing the multi-objective PPO path planning algorithm through a defined dynamic reward function and policy update mechanism, where the formula of the dynamic reward function is: ; in the formula, R is the total reward, , , are weight coefficients used to balance the importance of each reward term, is the path completion degree, indicating the completion degree of the current path, with a value range between 0 and 1, is the energy consumption, indicating the energy consumed by the manipulator structure in the operation task, and the calculation method is the sum of the squares of the joint torques. The specific formula is: , is the torque of the th joint, is the total number of joints, It represents the collision risk, which is the reciprocal of the distance between the robotic arm structure and the nearest obstacle and is normalized. A policy network structure is constructed and trained. The policy network structure consists of an LSTM layer and a fully connected layer for outputting the probability distribution of each joint movement. During the training process of the policy network structure, the state encoding result is input, the probability distribution of each joint movement in the corresponding environmental state is output, and the total reward in the corresponding environmental state is calculated by combining the dynamic reward function. The parameters of the policy network structure are updated based on the total reward in the corresponding environmental state to obtain an optimized multi-objective PPO path planning algorithm. The dynamic obstacles in the environment are detected in real time by a binocular vision camera, and according to the corresponding state encoding result, based on the defined local path update time, the local path replanning mechanism is triggered to replan the local path. Based on the optimized multi-objective PPO path planning algorithm, the probability distribution of each joint movement in the corresponding local path is output, and the optimal joint movement is selected. The optimal joint movement is used as the real-time obstacle avoidance path. In response to the real-time obstacle avoidance path, the path planning task of the corresponding real-time obstacle avoidance path is executed.

[0047] Furthermore, environmental perception and state encoding are key steps in the path planning of flying robots. The purpose is to extract useful environmental information from sensor data and encode it into a state representation that can be understood by machines. In this embodiment, the real-time environmental perception data includes real-time point cloud data collected by a binocular vision camera (such as a binocular vision camera of model Intel RealSense D435) and IMU attitude data collected by an IMU attitude sensor. Among them, the real-time point cloud data contains the three-dimensional spatial information of the objects in the environment. However, the original point cloud data is usually very large, and directly processing it will consume a large amount of computing resources. Therefore, it is first necessary to perform downsampling processing on the point cloud data and convert it into a depth map of 128×128 pixels. The depth map is a two-dimensional image, where each pixel value represents the distance between the object corresponding to the pixel and the binocular vision camera; through downsampling processing, the data volume can be significantly reduced while retaining key environmental information. Subsequently, the downsampled depth map is input into a three-layer convolutional neural network (CNN) for feature extraction; CNN is a deep learning model specifically used to process image data, and its core idea is to extract useful features from images through convolutional operations. In this embodiment, the three-layer structure of the CNN performs different feature extraction tasks. For example, the first convolutional layer: uses 32 3×3 convolutional kernels to perform convolutional operations on the input depth map and extract local features. The convolutional kernel slides on the image, and the feature value of a local area is calculated each time. Through the first convolutional layer, the edge and contour information in the depth map can be captured; the second convolutional layer: uses 64 3×3 convolutional kernels to further convolve the feature map output by the first layer. The purpose of this layer is to extract more advanced features, such as the shape and texture information of the object; the third convolutional layer: uses 128 3×3 convolutional kernels to perform the final convolutional operation on the feature map output by the second layer. The features extracted by this layer will be used for subsequent path planning decisions. After each convolutional operation, a non-linear transformation is performed through the ReLU activation function to enhance the expression ability of the model. In addition, max pooling operations are performed after each convolution to further reduce the size of the feature map while retaining the most important feature information. Through the convolutional operations of these three layers of CNN, the finally obtained feature map can effectively represent the key information in the environment, such as the position and shape of obstacles and the distance relationship between the robot and the obstacles; these features will be encoded into a state representation that can be understood by machines and used as the input for subsequent path planning strategies.

[0048] Furthermore, in the path planning of flying robots, policy optimization is a crucial step to ensure that flying robots can efficiently and safely complete tasks in complex environments. This application adopts a multi-objective PPO path planning algorithm to optimize the path planning strategy. Specifically, the multi-objective PPO path planning algorithm makes targeted adjustments in the design of the dynamic reward function and the policy update mechanism. Among them, the design of the dynamic reward function directly affects the behavior of the flying robot in path planning. In this embodiment, the dynamic reward function consists of three reward terms, namely the path completion degree , energy consumption , and collision risk . The specific formula is: ; where R is the total reward, , , are weight coefficients used to balance the importance of each reward term. is the path completion degree, indicating the completion degree of the current path. The value range is between 0 and 1. The higher the path completion degree, the greater the reward. The specific calculation formula is: = the length of the path completed by the flying robot / the total path length; is the energy consumption, indicating the energy consumed by the manipulator structure in the operation task. The calculation method is the sum of the squares of the joint torques. The lower the energy consumption, the greater the reward. The specific formula is: , is the torque of the th joint, is the total number of joints; is the collision risk, indicating the reciprocal of the distance between the manipulator structure and the nearest obstacle and is normalized. The lower the collision risk, the greater the reward. The specific formula is: = 1 / the distance between the manipulator structure and the nearest obstacle.

[0049] Furthermore, in this application, a policy network structure composed of an LSTM layer (long short-term memory network) and a fully connected layer is adopted to output the probability distribution of the motion of each joint. Among them, LSTM can capture long-term dependencies in time series data and is suitable for path planning in dynamic environments; the fully connected layer is used to output the probability distribution of the motion of each joint. The specific structure of the policy network structure includes: (1) LSTM layer: The input is the state encoding result , and the output is the hidden state . The calculation formula is: . In the formula, is the hidden state of the previous moment; (2) Fully connected layer: The output of the LSTM layer Map to the probability distribution of each joint movement through a fully connected layer; in a dynamic environment, a flying robot needs to detect dynamic obstacles in real time and quickly update the local path when an obstacle is detected. The specific implementation process is as follows: Detect dynamic obstacles in the environment in real time through a binocular vision camera, and based on the corresponding state encoding results, trigger the local path replanning mechanism based on the defined local path update time (such as 50 ms), replan the local path, and based on the optimized multi-objective PPO path planning algorithm, output the probability distribution of each joint movement under the corresponding local path, select the optimal joint movement, and use the optimal joint movement as the real-time obstacle avoidance path. In response to the real-time obstacle avoidance path, the flying robot executes the path planning task of the corresponding real-time obstacle avoidance path.

[0050] Furthermore, in order to trigger local path updates within the defined local path update time (such as 50 ms), this embodiment can also optimize the replanning mechanism. For example, (1) Parallel computing: Parallelize the dynamic obstacle detection and path replanning tasks, and use a multi-core CPU or GPU to accelerate the calculation; (2) Path caching: Pre-compute multiple possible local paths and store them in the path cache; when a dynamic obstacle is detected, directly select a suitable path from the cache to reduce the calculation time; (3) Incremental update: When updating the local path, only update the part of the path affected by the dynamic obstacle, rather than replanning the entire path, further reducing the calculation amount.

[0051] Embodiment 2

[0052] Based on the above Embodiment 1, this embodiment provides an adaptive disturbance rejection control system for a six-degree-of-freedom manipulator for high-altitude operations. Refer to Appendix Figure 4 , an adaptive disturbance rejection control method for a six-degree-of-freedom manipulator for high-altitude operations in Embodiment 1. The system includes a structure manufacturing module, a real-time dynamic disturbance rejection compensation module, an impedance parameter adjustment module, and an obstacle avoidance path planning module.

[0053] Furthermore, the structure manufacturing module is used to manufacture the manipulator structure by using an arm body with a carbon fiber topology optimization design and a joint housing provided with a flow guiding groove.

[0054] Furthermore, the real-time dynamic disturbance rejection compensation module is used to construct a dynamic disturbance observer based on the DDPG algorithm, output a dynamic compensation amount through the dynamic disturbance observer, combine the disturbance compensation algorithm, calculate the wind disturbance value and integrate it into the control law of the unmanned aerial vehicle for real-time dynamic disturbance rejection compensation.

[0055] Furthermore, the impedance parameter adjustment module is used to online adjust the impedance parameters based on the TD3 algorithm to cope with sudden changes in the load.

[0056] Further, the obstacle avoidance path planning module is used to generate a real-time obstacle avoidance path based on the optimized multi-objective PPO path planning algorithm and execute the path planning task of the corresponding real-time obstacle avoidance path to deal with dynamic obstacles.

[0057] Embodiment III

[0058] Based on the above Embodiment I, this embodiment provides a six-degree-of-freedom robotic arm for high-altitude operations, referring to the appendix Figures 5 - 6 , and the robotic arm structure is made by using the adaptive disturbance rejection control method of the six-degree-of-freedom robotic arm for high-altitude operations in Embodiment I.

[0059] Further, the robotic arm structure is composed of an arm body with a carbon fiber topology optimization design and a joint housing provided with a flow guiding groove.

[0060] Further, referring to the appendix Figure 6 , the arm body with a carbon fiber topology optimization design is an internal truss structure generated by performing multi-condition topology optimization on the arm body based on finite element analysis software, and the internal truss structure is made of carbon fiber composite materials. Among them, the load conditions for multi-condition topology optimization include a wind pressure of 15 m / s and an end load of 10 kg, the constraint conditions are a maximum stress ≤ 300 MPa and a first-order modal frequency ≥ 50 Hz, and the objective function includes minimizing the mass of the arm body and maximizing the stiffness of the arm body.

[0061] Further, referring to the appendix Figure 5 , the flow guiding groove is arranged on the surface of the joint housing and has a V-shaped structure. Geometric parameters including groove depth, inclination angle, and spacing are defined, and the wind resistance coefficient under the wind pressure condition of 15 m / s is simulated and verified through CFD software to optimize the aerodynamic performance.

[0062] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0063] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more flows and / or blocks in the process Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or blocks.

[0064] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one or more flows and / or blocks in the process Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or blocks.

[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks in the process Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or blocks.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. Adaptive disturbance rejection control method for a six-degree-of-freedom manipulator for high-altitude operations, characterized in that, Including: An arm body with a carbon fiber topology optimization design and a joint housing with a flow guiding groove are used to fabricate the manipulator structure; A dynamic disturbance observer is constructed based on the DDPG algorithm. The dynamic compensation amount is output through the dynamic disturbance observer. Combining with the disturbance compensation algorithm, the wind disturbance value is calculated and integrated into the control law of the unmanned aerial vehicle for real-time dynamic disturbance resistance compensation; Based on the TD3 algorithm, the impedance parameters are adjusted online to cope with sudden changes in load; Based on the optimized multi-objective PPO path planning algorithm, a real-time obstacle avoidance path is generated, and the path planning task of the corresponding real-time obstacle avoidance path is executed to cope with dynamic obstacles.

2. The adaptive disturbance rejection control method for the six-degree-of-freedom manipulator for high-altitude work according to claim 1, characterized in that, The arm body with the carbon fiber topology optimization design is an internal truss structure generated by performing multi-condition topology optimization on the arm body based on finite element analysis software. The internal truss structure is made of carbon fiber composite materials. Among them, the load conditions for multi-condition topology optimization include a wind pressure of 15 m / s and an end load of 10 kg. The constraint conditions are that the maximum stress ≤ 300 MPa and the first-order modal frequency ≥ 50 Hz. The objective function includes minimizing the mass of the arm body and maximizing the stiffness of the arm body.

3. The adaptive disturbance rejection control method for the six-degree-of-freedom manipulator for high-altitude operations according to claim 2, wherein, The flow guiding groove is arranged on the surface of the joint housing and has a V-shaped structure. Geometric parameters including groove depth, inclination angle, and spacing are defined. The wind resistance coefficient under the wind pressure condition of 15 m / s is simulated and verified by CFD software to optimize the aerodynamic performance.

4. The adaptive disturbance rejection control method for the six-degree-of-freedom manipulator for high-altitude operations according to claim 1, characterized in that, The construction of the dynamic disturbance observer based on the DDPG algorithm, the output of the dynamic compensation amount through the dynamic disturbance observer, and the combination with the disturbance compensation algorithm to calculate the wind disturbance value and integrate it into the control law of the unmanned aerial vehicle include: Collect real-time wind pressure data, joint state data, and historical error data to construct a state space data set; Construct and train the DDPG algorithm, and construct a dynamic disturbance observer based on the trained DDPG algorithm. Among them, the DDPG algorithm consists of a first Actor network for outputting the dynamic compensation amount and a first Critic network for evaluating the value of the dynamic compensation amount. Both the first Actor network and the first Critic network are composed of a three-layer fully connected structure, and the output nodes of each layer of the fully connected structure are 256, 128, and 64 respectively. During the training process of the DDPG algorithm, the state space data set is input, the dynamic compensation amount in the corresponding state space is output, the loss value in the real-time state space is calculated according to the mean square tracking error loss function, and the parameters of the DDPG algorithm are updated online in combination with the real-time wind pressure data to obtain the online updated DDPG algorithm; Based on the online updated DDPG algorithm, the dynamic compensation amount in the real-time state space is output, combined with the disturbance compensation algorithm, the wind disturbance value is calculated and integrated into the control law of the unmanned aerial vehicle for real-time dynamic disturbance resistance compensation.

5. The adaptive disturbance rejection control method for the six-degree-of-freedom robotic arm for working at height according to claim 4, characterized in that, The real-time wind pressure data is collected by multiple pressure sensors arranged on the surface of the manipulator structure. The joint state data includes the joint angle, joint speed, and current loop torque feedback by the joint encoder. The historical error data is the end pose tracking error data within the defined historical control period.

6. The adaptive disturbance rejection control method for the six-degree-of-freedom manipulator for high-altitude work according to claim 4, characterized in that, The formula of the disturbance compensation algorithm is: In the formula, is the wind disturbance value, is the dynamic compensation amount, is the state space data set, and are obtained after calibration by the particle swarm optimization algorithm, where takes the value of 12, takes the value of 0.5, is the control error.

7. The adaptive disturbance rejection control method for the six-degree-of-freedom manipulator for high-altitude work according to claim 1, characterized in that, The online adjustment of the impedance parameters based on the TD3 algorithm includes: Collect real-time end contact force, joint torque deviation data, and trajectory error data to construct an observation state dataset. Among them, the end contact force includes three force components and three torque components collected by a six-axis force sensor. The three force components are Fx, Fy, and Fz respectively, and the three torque components are Mx, My, and Mz respectively. The joint torque deviation data is the difference between the actual torque of each joint and the expected value. The trajectory error data includes the end position deviation value and the end velocity deviation value. Build and train the TD3 algorithm. Among them, the TD3 algorithm consists of a second Actor network for outputting stiffness and damping and a second Critic network for preventing overestimation. The second Actor network is composed of a four-layer fully connected structure. The output nodes of each layer of the fully connected structure are 512, 256, 128, and 64 respectively. The second Critic network adopts a double Q-network structure. When training the TD3 algorithm, input the observation state dataset and combine the defined OU noise to obtain the trained TD3 algorithm. Based on the trained TD3 algorithm, output the stiffness and damping under the real-time observation state and use them as the impedance parameters under the real-time observation state. Based on the impedance parameters under the real-time observation state, perform PID control and execute the adjustment of the corresponding impedance parameters to cope with sudden load changes.

8. The adaptive disturbance rejection control method of the six-degree-of-freedom manipulator for high-altitude operations according to claim 1, characterized in that, The optimized multi-objective PPO path planning algorithm generates a real-time obstacle avoidance path and executes the path planning task of the corresponding real-time obstacle avoidance path, including: Collect real-time environmental perception data. The real-time environmental perception data includes real-time point cloud data collected by a binocular vision camera and IMU attitude data collected by an IMU attitude sensor. Downsample the real-time point cloud data to obtain a depth map of 128×128 pixels, and input it into a three-layer convolutional neural network for feature extraction to output dynamic obstacle features. Then, perform state encoding on the dynamic obstacle features and the IMU attitude data to obtain the state encoding result. Optimize the multi-objective PPO path planning algorithm through the defined dynamic reward function and policy update mechanism. Among them, the formula of the dynamic reward function is: Wherein, R is the total reward, , , are weight coefficients used to balance the importance of each reward item, is the path completion degree, indicating the completion degree of the current path, and its value range is between 0 and 1, is the energy consumption, representing the energy consumed by the robotic arm structure during the operation task, and the calculation method is the sum of the squares of the joint torques. The specific formula is: , is the -th joint torque, is the total number of joints, is the collision risk, representing the reciprocal of the distance between the robotic arm structure and the nearest obstacle, and is normalized; Build and train the policy network structure. The policy network structure consists of an LSTM layer and a fully connected layer for outputting the probability distribution of each joint movement. During the training process of the policy network structure, input the state encoding result, output the probability distribution of each joint movement in the corresponding environmental state, and calculate the total reward in the corresponding environmental state in combination with the dynamic reward function. Update the parameters of the policy network structure based on the total reward in the corresponding environmental state to obtain the optimized multi-objective PPO path planning algorithm. Real-time detect dynamic obstacles in the environment through a binocular vision camera, and based on the corresponding state encoding result, trigger the local path replanning mechanism based on the defined local path update time, replan the local path, output the probability distribution of each joint movement under the corresponding local path based on the optimized multi-objective PPO path planning algorithm, and select the optimal joint movement, and use the optimal joint movement as the real-time obstacle avoidance path. In response to the real-time obstacle avoidance path, execute the path planning task of the corresponding real-time obstacle avoidance path.

9. Adaptive disturbance rejection control system for a six-degree-of-freedom robotic arm for high-altitude operations, characterized in that, Include: A structure manufacturing module, which is used to manufacture a robotic arm structure with an arm body designed by carbon fiber topology optimization and a joint housing provided with a diversion groove; A real-time dynamic disturbance rejection compensation module, which is used to construct a dynamic disturbance observer based on the DDPG algorithm, output a dynamic compensation amount through the dynamic disturbance observer, combine a disturbance compensation algorithm, calculate a wind disturbance value and integrate it into the control law of the unmanned aerial vehicle for real-time dynamic disturbance rejection compensation; An impedance parameter adjustment module, which is used to online adjust impedance parameters based on the TD3 algorithm to cope with sudden load changes; An obstacle avoidance path planning module, which is used to generate a real-time obstacle avoidance path based on an optimized multi-objective PPO path planning algorithm and execute the path planning task of the corresponding real-time obstacle avoidance path to cope with dynamic obstacles.

10. Six-degree-of-freedom robotic arm for working at height, characterized in that, It includes a robotic arm structure manufactured by using the adaptive disturbance rejection control method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Active-disturbance-rejection controller design method and device, and storage medium

    CN115903510A

  • Four-rotor unmanned aerial vehicle active-disturbance-rejection control method based on DDPG

    CN118795918A

  • Mobile robot real-time obstacle avoidance double-layer path planning method in nuclear environment

    CN119668264A

  • Multi-robot trajectory planning method

    WO2022241808A1

Cited By

  • Mechanical arm posture stable adjusting method based on vibration feedback

    CN120491443A

  • Lightweight multi-mode tea garden picking robot control system

    CN120839791A

  • Mechanical arm servo system anti-interference control system based on disturbance observer

    CN121300045A