Motion planning method and device for tower crane in complex environment

Through the BiRRT* algorithm of differential flat analysis and direction bias combined with the NURBS curve reparameterization method, the problem of positioning, obstacle avoidance and pendulum reduction in tower cranes in complex environments is solved, and efficient comprehensive control effect is achieved, reducing lifting time and energy consumption, and reducing load swing.

CN120288657APending Publication Date: 2025-07-11BEIJING WUZI UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510711072.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to achieve comprehensive control of precise positioning, load obstacle avoidance and pendulum reduction in tower cranes in complex environments. Especially in the presence of obstacles, the control is difficult and the existing methods are less flexible, making it difficult to meet the requirements of optimality and trajectory shape.

Method used

The BiRRT* algorithm (DB-BiRRT) based on differential flat analysis and direction bias is used to combine the non-uniform rational B-spline (NURBS) curve reparameterization method. Through path planning and trajectory fitting, the lifting trajectory is optimized to achieve load obstacle avoidance and pendulum reduction. The non-dominant sorting genetic algorithm (NSGA-II) is used for multi-objective optimization, and combined with the DDPG reinforcement learning algorithm to improve the efficiency and quality of path planning.

Benefits of technology

The tower crane is realized in a comprehensive control of precise positioning, load obstacle avoidance and pendulum reduction in complex environments, improving the efficiency of path planning and trajectory optimization effect, reducing lifting time and energy consumption, and reducing residual swing of the load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120288657A_ABST
    Figure CN120288657A_ABST
Patent Text Reader

Abstract

The invention discloses a tower crane motion planning method and device in a complex environment, and the method comprises the steps: building a nonlinear model for comprehensively describing the motion characteristics and transient behaviors of a crane-load system on the basis of kinetic analysis, and carrying out differential flat analysis; a BiRRT * algorithm based on direction bias is provided, a node expansion process is optimized by introducing a target bias mechanism of fusion region probability sampling and based on an improved potential field function direction guiding mechanism, and the efficiency and quality of path planning are improved. An improved random tree extension mechanism is combined with a depth deterministic policy gradient (DDPG) of reinforcement learning, and adaptive adjustment of sampling direction and step length parameters is realized. Path points obtained through path planning serve as profile value points of trajectory planning, system full-state constraint conditions such as load obstacle avoidance and shimmy reduction are fully considered, multi-target trajectory planning is carried out based on an NURBS curve, and a multi-target comprehensive optimal trajectory on the aspects of total hoisting time, operation energy consumption and load shimmy suppression is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence technology and robots, and particularly relates to a tower crane motion planning method in a complex environment, and a tower crane motion planning device in a complex environment. Background Art

[0002] A tower crane (i.e., a tower crane system) is an important transportation and lifting tool, which is widely used for cargo lifting operations at construction sites such as docks and construction sites. During the lifting process, the cargo often swings and oscillates due to the action of inertial force, which not only reduces the work efficiency, but also easily causes accidents and poses great safety hazards. In addition, in actual cargo lifting operations, there may be obstacles such as buildings and stacked goods in the working area of the tower crane. How to design appropriate planning and control strategies to avoid obstacles while achieving autonomous and accurate positioning of the crane and stable control of the load is a quite challenging problem.

[0003] At present, crane automatic control algorithms can be divided into two categories: open-loop control and closed-loop control. A large number of studies have focused on bridge formula cranes, and relatively few studies have been conducted on tower cranes. Due to the rotary motion of the boom of the tower crane, its dynamic characteristics are more complex than those of bridge formula cranes and gantry cranes, showing the characteristics of underactuation, multi-input multi-output, nonlinearity, and strong coupling. Therefore, the control difficulty is greater. For the open-loop control of tower cranes, some scholars have introduced input shaping control for load swing reduction, making the load residual vibration zero by convolving the input signal with a modulation pulse. However, the common feature of these methods is poor flexibility and it is difficult to meet the requirements of optimality and changing the trajectory shape. There are also proposed methods based on trajectory planning, which achieve the positioning and load swing reduction of the crane, but none of them consider the obstacle avoidance problem in the actual environment. For closed-loop control, the existing technologies have designed corresponding feedback controllers for the positioning and swing reduction problems. However, these methods require accurate real-time values of the swing angle, which will require high sensor accuracy in actual situations and do not have the obstacle avoidance function in complex environments. For the obstacle avoidance problem of the load, some scholars have proposed a graph planning algorithm based on the S curve to obtain the best trajectory sequence with the transition time as the index, but only consider the smoothness of the obstacle avoidance path and trajectory, and do not consider problems such as swing angle suppression in combination with the characteristics of tower cranes.

[0004] Although there has been some research in the field of tower crane automatic control, in the actual complex working environment, how to comprehensively and efficiently solve problems such as positioning, obstacle avoidance, and swing reduction is still an open problem to be improved. Summary of the Invention

[0005] In order to overcome the defects of the prior art, the technical problem to be solved by the present invention is to provide a tower crane motion planning method in a complex environment, which can independently complete the comprehensive control tasks of crane precise positioning, load obstacle avoidance and swing reduction, and has better performance in comparison.

[0006] The technical solution of the present invention is: the tower crane motion planning method in such a complex environment comprises the following steps:

[0007] (1) Based on differential flatness analysis, the motion planning problem of the system in a complex obstacle space is transformed into a planning problem with flat output;

[0008] (2) Based on the direction bias, the BiRRT* algorithm DB-BiRRT generates a reference path and introduces a direction guidance mechanism to improve the efficiency and quality of path planning;

[0009] (3) Based on the re-parameterization of the non-uniform rational B-spline NURBS curve, the fitting trajectory is optimized. Under the full-state constraints of the system, including load sway reduction and obstacle avoidance, a multi-objective optimization problem is proposed with the goal of minimizing the lifting time, energy consumption and sway. The shape parameters and time node distribution of the NURBS curve are optimized to obtain the comprehensive optimal trajectory curve of the system state. The non-dominated sorting genetic algorithm NSGA-II is used for optimization and solution.

[0010] The step (2) comprises the following sub-steps:

[0011] (2.1) The BiRRT algorithm based on direction bias guides the node generation process by introducing the target bias mechanism of fusion regional probability sampling and the direction guidance mechanism of improved potential field function, so that new nodes are generated in a more optimal direction;

[0012] (2.2) Construct a parameter decision link based on the DDPG reinforcement learning algorithm, dynamically generate important parameters in the proposed random tree node expansion process, thereby constructing an adaptive bidirectional RRT algorithm based on direction bias and designing a reward function;

[0013] (2.3) When selecting the path points of the tower crane system load, a greedy algorithm is used to eliminate redundant points.

[0014] Only critical path points are retained.

[0015] Based on the dynamic analysis, the present invention establishes a non-linear model that comprehensively describes the motion characteristics and transient behavior of the crane-load system, and conducts differential flatness analysis to provide a simple and direct expression formula for motion planning. Aiming at the path planning problem of tower cranes in complex environments, a BiRRT* algorithm based on direction offset is proposed. By introducing a target offset mechanism that fuses regional probability sampling and a direction guidance mechanism based on an improved potential field function to optimize the node expansion process, the efficiency and quality of path planning are improved. Then, the improved random tree expansion mechanism is combined with the deep deterministic policy gradient DDPG of reinforcement learning to achieve the adaptive adjustment of the sampling direction and step size parameters, effectively solving the defects of slow convergence speed and sub-optimal paths in the RRT series algorithms. After that, the path points obtained from path planning are used as the control points of trajectory planning. Under the full state constraints of the system such as load obstacle avoidance and swing reduction, a multi-objective optimization method based on NURBS curves is used for trajectory planning to obtain a multi-objective comprehensive optimal trajectory in terms of total lifting time, operating energy consumption, and load swing suppression. Therefore, the present invention can complete the comprehensive control tasks of accurate positioning of the crane, load obstacle avoidance, and swing reduction, and has better performance compared with others.

[0016] A motion planning device for tower cranes in complex environments is also provided, which includes:

[0017] A signal acquisition module and device, with a signal acquisition module provided at the trolley for obstacle information acquisition and data transmission;

[0018] A differential flatness analysis module, which based on differential flatness analysis, transforms the motion planning problem of the system in the complex obstacle space into a planning problem of flat output;

[0019] A path planning module, based on the obstacle information obtained by the signal acquisition module and device, generates a reference path based on the direction-biased BiRRT* algorithm DB-BiRRT, and introduces a direction guidance mechanism to improve the efficiency and quality of path planning;

[0020] A trajectory planning module, which based on the fitting of the lifting trajectory by reparameterizing the non-uniform rational B-spline NURBS curve, optimizes the fitted trajectory. Under the full state constraints of the system considering load swing reduction and obstacle avoidance, with the goal of minimizing the lifting time, energy consumption, and swing, a multi-objective optimization problem is proposed to optimize the shape parameters and time node distribution of the NURBS curve, obtain the comprehensive optimal trajectory curve of the system state, and use the non-dominated sorting genetic algorithm NSGA-II for optimization and solution;

[0021] The path planning module executes:

[0022] (2.1) The BiRRT algorithm based on direction bias guides the node generation process by introducing an objective bias mechanism that fuses regional probability sampling and a direction guidance mechanism that improves the potential field function, enabling new nodes to be generated in a more optimal direction.

[0023] (2.2) Construct a parameter decision link based on the DDPG reinforcement learning algorithm to dynamically generate important parameters in the proposed random tree node expansion process, thereby constructing an adaptive bidirectional RRT algorithm based on direction bias and designing a reward function.

[0024] (2.3) When selecting path points for the tower crane system load, the greedy algorithm is used to eliminate redundant points and only key path points are retained. Description of the Drawings

[0025] Figure 1 It is a schematic diagram of the physical structure of a tower crane.

[0026] Figure 2 It shows the physical meaning of the auxiliary state variables.

[0027] Figure 3 It shows the regional probability sampling and the resultant force on the node.

[0028] Figure 4 It shows the basic structure of DB-AD-BiRRT*.

[0029] Figure 5 It is the complete flowchart of DB-AD-BiRRT.

[0030] Figure 6 It shows the principle of node elimination.

[0031] Figure 7 It shows the training reward curve.

[0032] Figure 8 It shows the path planning results of four algorithms.

[0033] Figure 9 It shows the running times of four algorithms.

[0034] Figure 10 It shows the path lengths of four algorithms.

[0035] Figure 11 It shows the number of nodes of four algorithms.

[0036] Figure 12 It shows the obstacle avoidance trajectory of the tower crane system load.

[0037] Figure 13 It shows the displacement response curve of the tower crane (Experiment 1).

[0038] Figure 14Shows the speed response curve of the tower crane (Experiment 1).

[0039] Figure 15 Shows the acceleration response curve of the tower crane (Experiment 1).

[0040] Figure 16 Shows the displacement response curve of the tower crane (Experiment 2).

[0041] Figure 17 Shows the speed response curve of the tower crane (Experiment 2).

[0042] Figure 18 Shows the acceleration response curve of the tower crane (Experiment 2). Detailed implementation method

[0044] The motion planning method of the tower crane in such a complex environment includes the following steps:

[0045] (1) Based on differential flatness analysis, transform the motion planning problem of the system in the complex obstacle space into a planning problem of the flat output;

[0046] (2) Based on the direction-biased BiRRT* algorithm DB-BiRRT, generate a benchmark reference path, and introduce a direction guidance mechanism to improve the efficiency and quality of path planning;

[0047] (3) Based on the reparameterization of the non-uniform rational B-spline NURBS curve for the hoisting trajectory fitting, optimize the fitted trajectory, and considering the system full-state constraints of load swing reduction and obstacle avoidance

[0048] Under this condition, with the goal of minimizing the hoisting time, energy consumption, and swing, propose a multi-objective optimization problem to optimize the shape parameters and time node distribution of the NURBS curve, obtain the comprehensive optimal trajectory curve of the system state, and use the non-dominated sorting genetic algorithm NSGA-II for optimization and solution;

[0050] The step (2) includes the following sub-steps:

[0051] (2.1) Based on the direction-biased BiRRT algorithm, guide the node generation process by introducing a target biasing mechanism that fuses regional probability sampling and a direction guidance mechanism that improves the potential field function, so that

[0052] The new nodes are generated in a more optimal direction;

[0053] (2.2) Construct a parameter decision-making link based on the DDPG reinforcement learning algorithm, dynamically generate important parameters in the random tree node expansion process proposed, so as to construct a self

[0054] Adaptive bidirectional RRT algorithm and design a reward function;

[0055] (2.3) When selecting the path points of the tower crane system load, the greedy algorithm is used to eliminate redundant points and only key path points are retained.

[0056] Based on the dynamic analysis, the present invention establishes a nonlinear model that comprehensively describes the motion characteristics and transient behavior of the crane-load system, and conducts differential flatness analysis to provide a simple and direct expression formula for motion planning. Aiming at the path planning problem of tower cranes in complex environments, a BiRRT* algorithm based on direction offset is proposed. By introducing an objective offset mechanism that fuses regional probability sampling and a direction guidance mechanism based on an improved potential field function to optimize the node expansion process, the efficiency and quality of path planning are improved. Then, the improved random tree expansion mechanism is combined with the deep deterministic policy gradient DDPG of reinforcement learning to realize the adaptive adjustment of the sampling direction and step size parameters, effectively solving the defects of slow convergence speed and sub-optimal paths in the RRT series algorithms. After that, the path points obtained from path planning are used as the shape value points for trajectory planning. Under the full consideration of system full-state constraints such as load obstacle avoidance and swing reduction, a multi-objective optimization method based on NURBS curves is used for trajectory planning to obtain a multi-objective comprehensive optimal trajectory in terms of total lifting time, operating energy consumption, and load swing suppression. Therefore, the present invention can complete the comprehensive control tasks of accurate positioning of the crane, load obstacle avoidance, and swing reduction, and has better performance compared with others.

[0057] Preferably, in the step (1), according to formula (16), all state variables of the system are represented by (X a , Y a , x, l) and their differential variables. The system is differentially flat. For the motion planning of the tower crane system, it is equivalently transformed into the motion planning of the flat output (X a , Y a , x, l).

[0058]

[0059] Among them, is the conversion relationship between the states θ1, θ2, θ3 and the flat output.

[0060] Preferably, in the step (2.1), when expanding the random tree, there is a probability of directly expanding towards the target point, thereby accelerating the expansion speed towards the target point:

[0061]

[0062] Among them, q rand , q goal , q pro are the random sampling point, the target point, and the regional probability sampling point, P0 is the bias probability, and P ∈ (0, 1) is a random number;

[0063] When generating sampling nodes without performing target biasing, the entire map is evenly divided into countless random sampling regions. Based on the line connecting the starting point and the target point, the distance from each random sampling region to this line is used as a variable to calculate the sampling probability of each region, and the Gaussian distribution is used to determine the sampling probability in each region.

[0064]

[0065] Among them, d c is the distance from the center point of the region to the line connecting the starting and ending points, l is the distance between the starting and ending points, the standard deviation of the Gaussian distribution probability P(d c ) is 0.25l, p st , p end represent the starting point and the ending point respectively, and p c is the center point of these sampling regions.

[0066] Preferably, in the step (2.1), the potential field function is improved:

[0067]

[0068] Among them, α i (i = 1, 2), β are the attraction coefficient and the repulsion coefficient respectively, q, q o , q g , q r represent the current point, the obstacle, the target point and the sampling point respectively, d(q o ), d g (q g ), d r (q r ) are the Euclidean distances between the current point and the obstacle, the target point and the sampling point respectively, d a is the limit range of the maximum target attraction, d0 is the maximum distance at which the potential field takes effect, m is the repulsion intensity parameter, and it satisfies m e > 0;

[0069] The attraction and repulsion functions are the negative gradients of the potential field function. The attraction F att and the repulsion

[0070] force F req of the current node are:

[0071]

[0072]

[0073] Among them, the direction of F att1 is from the current node to the target point, the direction of F att2 is from the current node to the sampling point, and the direction of F reqThe direction points from the obstacle to the current node.

[0074] Preferably, in the step (2.2), the constructed random tree expansion mechanism is modeled as the environment of reinforcement learning, including the following:

[0075] (2.2.1) State space: To obtain sufficient environmental information of the nodes, the state space is designed as:

[0076] S = {D g1 , D g2 , D r1 , D r2 , D ob1 , D ob2 , N p1 , N p2} (23)

[0077] Wherein, D g1 , D g2 are respectively the 2D vector values from the current nodes of the two trees to the corresponding target points,

[0078] D r1 , D r2 are respectively the 2D vector values from the current nodes of the two trees to their nearest randomly sampled points,

[0079] D ob1 , D ob2 are the 2D vector values from the current nodes of the two trees to their nearest obstacles, N p1 , N p2 is the total number of nodes within the range O of the current nodes of the two trees, and the entire state space has a total of 14 dimensions;

[0080] (2.2.2) Action space: According to the design and parameter requirements of the improved random tree expansion mechanism, it is designed as:

[0081] A = {P1, P2, β1, β2, L1, L2} (24)

[0082] Wherein, P1 and P2 are respectively the bias probabilities of the two trees towards the target points, β1 and β2 are the current repulsive force coefficients of the two trees, and L1 and L2 are the current expansion step lengths of the two trees. The entire action space has a total of 6 dimensions;

[0083] (2.2.3) Reward function: Considering the sparse reward problem and the optimization effect comprehensively, the reward function design includes: the obstacle avoidance reward when the random tree expands nodes, the proximity reward when the random tree expands nodes, the connection reward after obtaining a feasible path, and the path optimization reward after obtaining a feasible path.

[0084] Preferably, in the step (2.2.3),

[0085] Obstacle avoidance reward: Each time an extended node collides with an obstacle or is less than the safe distance, a reward value of -1 is given, and the reward is 0 at other times, expressed as:

[0086]

[0087] where d ob is the distance to the nearest obstacle, and d safe is the set safe distance;

[0088] Approaching reward: When the generated node is far from the target point, a reward value of -1 is given, and the reward is 0 at other times, expressed as:

[0089]

[0090] where d goal is the distance from the current node to the target point, and d b is the distance from its parent node to the target point;

[0091] Connection reward: If two random trees are successfully connected to generate a feasible path, a reward of 100 is given, and if the connection fails, a reward value of -200 is given, expressed as:

[0092]

[0093] Path optimization reward: Reward is given for the length of the obtained path. If the path exceeds the limit value X lim a penalty is given. When the path is the current global optimum and does not exceed the limit, a reward is given, and the reward is 0 at other times, expressed as:

[0094]

[0095] Preferably, in the step (2.2), the training process of the decision-making model is as follows:

[0096] (2.2.a) Initialize the Actor-Critic network, set the initial weights of the network and initialize the experience replay pool, and reset the reinforcement learning environment, including the state space, action space, reward, and random tree nodes;

[0097] (2.2.b) The agent selects an action according to the current state and policy, obtains the next node according to the improved random tree expansion mechanism, and calculates the reward for this action according to (25)-(28). During the expansion process, record the current state, the action taken, the reward obtained, and the next state, and store them in the experience replay pool;

[0098] (2.2.c) Randomly select a certain number of samples from the experience replay pool, and update the weights of the actor and critic networks through the gradient descent method. In each update, the goal is to maximize the Q value output by the Critic network, thereby optimizing the agent's policy selection;

[0099] (2.2.d) Repeat the above training steps. As the training progresses, the agent's policy will be gradually improved and tend to be optimized. Eventually, the reward curve will tend to be stable, and the training will be completed.

[0100] Preferably, in step (3), optimize the fitted (X a , Y a , x, l) trajectory, and design the multi-objective optimization decision variable X as follows:

[0101]

[0102] In the formula, ω is the weight factor vector of the curve, and t is the vector of the time intervals of the shape value points. The number of elements in the two vectors is 2z + 14 and z - 1 respectively;

[0103] During the hoisting process, the state constraints of the operating room:

[0104]

[0105] Constrain the swing angle and its speed:

[0106]

[0107] In the formula, v ilim (i = 1,..., 5) is the maximum value constraint of the speed of each physical quantity, and a ilim (i = 1, 2, 3) is the maximum value constraint of the acceleration of the all-driven physical quantity, and θ lim is the maximum value constraint of the swing angle;

[0108] Let Ω be the set of non-obstacle positions, and constrain the generalized coordinates of the load:

[0109] (x, l, θ1, θ2, θ3) ∈ Ω (42)

[0110] Select 3 objective functions for the multi-objective optimization problem as:

[0111]

[0112] In the formula, f0(x) is the hoisting time, v x (t), v θ1 (t) and v l(t) are the luffing displacement, slewing angle, and rope length velocity respectively, f1(x) is the total energy consumption of the actuator, f2(x) is the evaluation function of the swing angle, x,y∈(0,1) are user-defined weights, and max|θ j |(j = 2,3) represents the maximum absolute value of the swing angle.

[0113] Preferably, in step (3), the state variables are replaced with the flat output form formula by the conversion relationship of formula (16), and the multi-objective optimization model of the (X a ,Y a ,x,l) trajectory is:

[0114]

[0115] where g1 to g5 are the constraints satisfied by the optimization problem;

[0116] To obtain solutions that meet different objective priority requirements, let the normalized weight objective function be formula (47), and select a trade-off optimal solution with the minimum of f op . Adjust the optimal solution according to different task requirements or user preferences by setting different weights λ1, λ2, λ3

[0117]

[0118] In the formula, n is the number of population individuals, f 0i ,f 1i ,f 2i are the objective function values of the i-th individual of the above three objective functions respectively, f 0min ,f 1min ,f 2min are the minimum values of the three objective functions among all individuals, f 0max ,f 1max ,f 2max are the maximum values of the three objective functions among all individuals, and λ1, λ2, λ3 are the weights of the three objective functions, which are set according to preferences or tasks as needed.

[0119] Those of ordinary skill in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes the steps of the above method embodiments, and the storage medium can be: ROM / RAM, magnetic disk, optical disk, memory card, etc. Therefore, corresponding to the method of the present invention, the present invention also simultaneously includes a tower crane motion planning device in a complex environment, and this device is usually represented in the form of functional modules corresponding to the steps of the method. This device includes:

[0120] Signal acquisition module and device, a signal acquisition module is provided at the trolley for obstacle information acquisition and data transmission;

[0121] Differential flatness analysis module, which, based on differential flatness analysis, transforms the motion planning problem of the system in a complex obstacle space into a planning problem of flat output;

[0122] Path planning module, which generates a reference path based on the direction-biased BiRRT* algorithm DB-BiRRT, and introduces a direction guidance mechanism to improve the efficiency and quality of path planning;

[0123] Trajectory planning module, which fits the hoisting trajectory based on the reparameterization of non-uniform rational B-spline (NURBS) curves, optimizes the fitted trajectory, and, considering the full-state constraints of the system such as load swing reduction and obstacle avoidance, aims to minimize the hoisting time, energy consumption, and swing, proposes a multi-objective optimization problem to optimize the shape parameters and time node distribution of the NURBS curve, obtains the comprehensive optimal trajectory curve of the system state, and uses the non-dominated sorting genetic algorithm NSGA-II for optimization and solution;

[0124] The path planning module performs:

[0125] (2.1) Based on the direction-biased BiRRT algorithm, by introducing a target biasing mechanism that fuses regional probability sampling and a direction guidance mechanism that improves the potential field function to guide the node generation process, enabling new nodes to be generated in a more optimal direction;

[0126] (2.2) Construct a parameter decision link based on the DDPG reinforcement learning algorithm to dynamically generate important parameters in the proposed random tree node expansion process, thereby constructing an adaptive bidirectional RRT algorithm based on direction bias and designing a reward function;

[0127] (2.3) When selecting path points for the load of the tower crane system, use the greedy algorithm to eliminate redundant points and only retain key path points.

[0128] The content of the present invention is described in more detail below.

[0129] 1 Tower crane dynamics model and flatness analysis

[0130] 1.1 Dynamics model

[0131] The hoisting operation of the tower crane mainly consists of three parts: lifting, slewing, and luffing. Establish its simple physical structure as Figure 1 shown.

[0132] s i =sin(θ i ),c i =cos(θ i ),i=1,2,3 (1)

[0133] Let \(M\) be the mass of the trolley, \(m\) be the mass of the load, \(x\) be the luffing displacement of the trolley, \(l\) be the length of the hoisting rope, \(\theta_1\) be the slewing angle of the boom, \(\theta_2,\theta_3\) be the swinging angles of the load, and \(F_1,\tau,F_2\) be the system inputs, representing the control force for the trolley displacement, the control torque for the boom slewing, and the control force for the hoisting rope length respectively. Let \(g\) be the acceleration due to gravity and \(J\) be the moment of inertia of the system. Then the total kinetic energy \(K\) of the system consists of the kinetic energies of the boom, the trolley, and the load:

[0134]

[0135] where \((X_1,Y_1,Z_1)\) and \((X_2,Y_2,Z_2)\) are the coordinates of the trolley and the load.

[0136] Taking the XOY plane as the zero potential energy plane of the system, the total potential energy \(P\) of the system consists of the potential energy of the load:

[0137] \(P = -lmgc_2c_3\ (3)\)

[0138] Construct the Lagrangian operator \(L = K - P\) of the system and list the Lagrangian dynamics equations:

[0139]

[0140] Establish a complete dynamic mathematical model of the system.

[0141] Since the swinging angles should be kept within a small range during the normal operation of the tower crane, the underactuated part of the dynamics equations can be approximated, that is, it is considered that:

[0142]

[0143] Obtain the simplified dynamics equations for the underactuated part of the system:

[0144]

[0145] 1.2 Flatness analysis

[0146] According to the dynamics equations, the tower crane with variable rope length has a total of five degrees of freedom \(x,\theta_1,l,\theta_2,\theta_3\) and three control input quantities \(F_1,\tau,F_2\), which is an underactuated system. There is a strong coupling relationship between the load swing and the trolley movement, and the dynamics of the system is relatively complex. To facilitate the handling of the coupling problem and simplify the planning process, the movements of the trolley and the load are correlated, and an auxiliary state quantity \((X Figure 2 shown in a ,Y a ) of the load position is constructed as follows:

[0147] \(X a = xc_1 + l\theta_3c_1 - l\theta_2s_1\ (8)

[0148] Y a = xs1 + lθ3s1 + lθ2c1 (9)

[0149] Furthermore, the following theorem is proposed:

[0150] Theorem 1: The hoisting system of a tower crane with variable rope length is a differentially flat system. The motion planning for this system can be transformed into the motion planning of the flat output (X a , Y a , x, l).

[0151] Proof:

[0152] When the load swing angles θ2 and θ3 are small, the load position can be approximated by equations (8)-(9). Multiply equation (8) by s1, equation (9) by c1, and subtract them. Multiply equation (8) by c1, equation (9) by s1, and add them. After arrangement, we can get:

[0153]

[0154] Multiply the second derivative of equation (8) by s1, the second derivative of equation (9) by c1, and then subtract them to get equation (12). Multiply the second derivative of equation (8) by c1, the second derivative of equation (9) by s1, and then add them to get equation (13):

[0155]

[0156] Combining equations (6)-(7) and equations (12)-(13), we can get:

[0157]

[0158] Substitute equations (10)-(11) into equations (14)-(15) and arrange them to obtain equation (16).

[0159] According to equation (16), all state variables of the system can be represented by (X a , Y a , x, l) and their differential variables. Also, according to the dynamic equations in the appendix, obviously the system inputs (F1, F2, τ) can be represented by the state variables. Therefore, the inputs can also be represented by (X a , Y a , x, l) and their differential variables. Thus, the system is differentially flat, and the motion planning for the tower crane system can be equivalently transformed into the motion planning of the flat output (X a , Y a , x, l).

[0160] 2 Tower Crane Path Planning Based on DB - BiRRT*

[0161] In a complex environment, there are many obstacle constraints. Path planning can search for a safe and feasible path from the starting point to the target point based on information such as the movement range of the crane, the distribution of obstacles, the starting point, and the target point. Path planning is the basic link in the entire motion planning process. It determines the reference positions of various states during the operation of the tower crane system and provides a key reference path for subsequent trajectory planning.

[0162] 2.1 BiRRT* Algorithm Based on Direction Biasing

[0163] The Rapidly exploring Random Tree (RRT) is an algorithm that quickly searches by randomly constructing a space-filling tree and is suitable for solving the path planning problem of tower cranes in complex environments. The BiRRT* algorithm improves the RRT algorithm by using the method of simultaneously growing and expanding paths from the starting point and the target point to find a better connection method for optimizing the path cost of mobile robots. However, in a narrow or complex space, it may still encounter the problem of being trapped in a local optimum and difficult to escape. At the same time, since the nodes are randomly generated during sampling, there are too many "branches" generated, with a large number of redundant nodes, and the search lacks directionality, resulting in low algorithm efficiency.

[0164] To address the above deficiencies, this paper proposes the BiRRT algorithm based on direction biasing (DB-BiRRT), which guides the node generation process by introducing a target biasing mechanism that fuses regional probability sampling and a direction guiding mechanism that improves the potential field function, enabling new nodes to be generated in a better direction.

[0165] The specific improvement mechanism of the algorithm is as follows:

[0166] (1) Target Biasing Mechanism with Fused Regional Probability Sampling

[0167] Target biasing mechanism: When expanding the random tree, there is a probability of directly expanding towards the target point, thus accelerating the expansion speed towards the target point, that is, Equation (17).

[0168] Regional probability sampling mechanism: When P > P0, that is, when the target biasing is not executed, a sampling method based on regional probability is proposed. When generating sampling nodes, the entire map is evenly divided into countless random sampling regions. Based on the line connecting the starting point and the target point, the distance from the random sampling region to this line is used as a variable to calculate the sampling probability of each region. The Gaussian distribution is used to determine the sampling probability in each region (Equation (18)).

[0169] (2) Search Mechanism Based on Direction Guidance

[0170] During the algorithm search process, the concepts of repulsive field and attractive field of the potential field function are introduced. In the traditional potential field function, the resultant force of repulsion and attraction may be zero, leading to the situation of falling into local optimum, or when the target point is close to the obstacle, the repulsion may continuously be greater than the attraction, thus unable to reach the target position. To enhance the reachability of the target point and prevent falling into local optimum, an improved potential field function is proposed, and the randomly sampled points after the above sampling are introduced into the potential field. When determining the expansion direction of the next tree, the attraction and repulsion of the current node are affected by three factors: the target point, the obstacle, and the randomly sampled point. The improved potential field function is as shown in equations (19)-(20).

[0171] The attraction and repulsion functions are the negative gradients of the potential field function. From equations (19)-(20), the attraction force F of the current node can be obtained att and the repulsion force F req , that is, equations (21) and (22).

[0172] The improved potential field function limits the attraction of the target point and reduces the repulsion when the path approaches the target point. Combining the attraction of the randomly sampled point, it thus has better target reachability and avoids falling into local optimum. The schematic diagrams of regional probability sampling and the resultant force F received by the node are as Figure 3 shown.

[0173] 2.2 Parameter decision-making link based on DDPG reinforcement learning algorithm

[0174] To make the algorithm have better adaptability and flexibility, a parameter decision-making link based on DDPG reinforcement learning algorithm is constructed to dynamically generate the important parameters in the expansion process of the proposed random tree nodes, thereby constructing an adaptive bidirectional RRT algorithm based on direction bias (DB-AD-BiRRT), and its basic structure is as Figures 3 - 5 shown:

[0175] The random tree expansion mechanism constructed in the previous section is modeled as the environment of reinforcement learning, which mainly includes the following parts:

[0176] (1) State space: To obtain sufficient environmental information of the node, the state space is designed as equation (23).

[0177] (2) Action space: According to the design and parameter requirements of the improved random tree expansion mechanism, it is designed as equation (24).

[0178] (3) Reward function: Considering the sparse reward problem and optimization effect comprehensively, the reward function design includes 4 parts, namely the obstacle avoidance reward, the approaching reward when the random tree expands nodes, the connection reward after obtaining a feasible path, and the path optimization reward.

[0179] Collision avoidance reward: Each time an extended node collides with an obstacle or is less than the safety distance, a reward value of -1 is given, and the reward is 0 at other times, which is expressed as Equation (25).

[0180] Approaching reward: When the generated node is far from the target point, a reward value of -1 is given, and the reward is 0 at other times, which is expressed as Equation (26).

[0181] Connection reward: If two random trees are successfully connected to generate a feasible path, a reward of 100 is given; if the connection fails, a reward value of -200 is given, which is expressed as Equation (27).

[0182] Path optimization reward: Reward is given for the length of the obtained path. If the path exceeds the limit value X lim a penalty is given. When the path is the current global optimum and does not exceed the limit, a reward is given, and the reward is 0 at other times, which is expressed as Equation (28).

[0183] The training process of the decision-making model is as follows:

[0184] Step 1: Initialize the Actor-Critic network, set the initial weights of the network, and initialize the experience replay pool. Reset the reinforcement learning environment, including the state space, action space, reward, and random tree nodes.

[0185] Step 2: The agent selects an action according to the current state and policy, obtains the next node (i.e., enters the next state) according to the improved random tree extension mechanism, and calculates the reward for this action according to Equations (25)-(28). During the extension process, record the current state, the action taken, the reward obtained, and the next state, and store them in the experience replay pool.

[0186] Step 3: Randomly sample a certain number of samples from the experience replay pool, and update the weights of the actor and critic networks by the gradient descent method. In each update, the goal is to maximize the Q value output by the Critic network, thereby optimizing the policy selection of the agent.

[0187] Step 4: Repeat the above training steps. As the training progresses, the policy of the agent will be gradually improved and tend to be optimized. Finally, the reward curve tends to be stable, and the training is completed.

[0188] The algorithm flow of the DB-AD-BiRRT algorithm is as Figure 5 shown. First, initialize parameters such as the starting and ending positions, obstacle map, and random trees. Then load the trained agent. The decision-making model makes corresponding actions according to the current environmental state, and then performs target biasing, regional probability sampling, and resultant force calculation according to the actions. Combining with the step size, a new sampling point is determined. After such cyclic sampling until the two random trees are successfully connected, the final feasible path is obtained, and the path planning is completed.

[0189] 2.3 Selection of Key Path Points

[0190] In path planning, redundant path points will greatly reduce the efficiency of subsequent trajectory planning. Therefore, when selecting the path points (X a * , Y a * ) of the tower crane system load, the greedy algorithm is used to eliminate redundant points and only key path points are retained. The steps are as follows:

[0191] Starting from the initial point, connect to each subsequent node. If there are no obstacles between the connections, remove all nodes between the two nodes and replace the original path with the connection between the two nodes. If there are obstacles between the connections, start this operation again from the parent node of this node until all nodes are traversed. The principle is as Figure 6 shown.

[0192] In addition, in order to obtain the path points of the trolley displacement x and the rope length l of the tower crane system, let the expected swing angle of the load be 0 when passing through the path point . Then the coordinates of the luffing trolley in the X and Y directions are also Therefore, according to the geometric relationship, the path points corresponding to x and l (x * , l * ) can be expressed by Equation (29).

[0193]

[0194] where l0, l d , x0, x d are the starting state and target state of the rope length and the luffing displacement respectively.

[0195] 3 Multi-objective Trajectory Planning Based on NURBS Curve

[0196] In the motion planning of the tower crane, the path planning generates a geometric obstacle avoidance path. The path points are sparse, tortuous and do not meet the kinematic and dynamic characteristics, and cannot be directly used in the control system. Therefore, on the basis of path planning, kinematic constraints such as speed and acceleration and dynamic constraints in time are further considered to generate an executable motion trajectory corresponding to the load, trolley displacement and rope length variables.

[0197] 3.1 Related Definitions of NURBS Curve

[0198] NURBS consists of three factors: control points, weight factors and knot vectors. A NURBS curve of degree p is defined as:

[0199]

[0200] where P iis the control point of the curve, ω i is the weight factor of the curve, and its value affects the local shape of the curve. N i,p N(u) is the p-th order B-spline basis function defined on the knot vector U, where i represents the serial number of the p-th order basis function.

[0201] The k-th order derivative of the NURBS curve can be expressed as:

[0202]

[0203] 3.2 Hoisting trajectory fitting based on NURBS curve reparameterization

[0204] In the flat space, according to the path planning, the type value points {Q i}, i = 1,..., z, are used to fit the flat output (X i , Y a , x, l) with a k a (i = 1, 2, 3, 4)-th order NURBS curve. Let the total hoisting time be T, and let the curve parameter u be the time variable t / T. Then the specific expression of the knot vector U is designed as Equation (33):

[0205]

[0206] where {t j} (j = 1,..., z - 1) is the time interval between z type value points.

[0207] According to the hoisting requirements, the tower crane state variables (x, l, θ1, θ2, θ3) need to start precisely from the initial state, stop at the end state, and have zero velocity and acceleration. Transformed to the flat space according to Equation (16), the derivative values of (X a , Y a , x, l) at the endpoints are expressed as Equation (34).

[0208]

[0209] Due to space limitations, taking the trajectory fitting of X a , Y a as an example, to satisfy the fourth-order derivative at the starting and ending points to be zero, boundary conditions D a = 0 (a = 1,..., 4) and D b = 0 (b = 1,..., 4) are added. Then the solution equations for the control points P are shown in Equations (35)-(38). In addition, to make the equations have a unique solution, the NURBS curve degrees k1 = k2 = 9, k3 = k4 = 7.

[0210] Q = NP (35)

[0211]

[0212] In the formula, D a = 0 4×1 , D b = 0 4×1 , which are the values of the profile points and the derivative at the endpoints. is the boundary derivative matrix, and the values of its matrix elements can be obtained from Equation (32) and the knot vector Equation (33). is the basis function matrix, and the matrix elements can be obtained from Equation (31) and the knot vector Equation (33).

[0213] From the above analysis, it can be seen that the elements in matrix N need to be calculated through the knot vector U. Therefore, the control points obtained by solving the above equations include the time interval {t j}. It can be obtained that the fitting NURBS hoisting trajectory includes the time interval {t j} (j = 1,..., z - 1) of the profile points and the weight factors {ω i} (i = 1,..., z + 8).

[0214] 3.3 Multi-objective optimization of the hoisting trajectory

[0215] Optimize the fitted (X a , Y a , x, l) trajectory. Considering that the weight factor of the NURBS curve only affects the local shape of the curve and does not affect the overall trajectory, in order to make the trajectory more flexible and have more optimization space, the multi-objective optimization decision variable X is designed as Equation (39).

[0216] During the hoisting process, the state constraints also need to be considered. Due to the limitations of the actuator, it is necessary to ensure that the speed and acceleration in the fully actuated state are not too large, thus obtaining the state constraints during operation, that is, Equation (40).

[0217] Considering the problem of reducing swing, it is necessary to constrain the swing angle and its speed, that is, Equation (41).

[0218] Considering the problem of obstacle avoidance, let Ω be the set of non-obstacle positions, and constrain the generalized coordinates of the load, that is, Equation (42).

[0219] The optimization objective is to balance among the hoisting time, energy consumption, and the magnitude of the swing angle. Therefore, the objective functions of the multi-objective optimization problem are selected as Equations (43)-(45).

[0220] According to the proven Theorem 1, it can be obtained that the constraints and objective functions of Equations (40)-(45) can all be expressed by the flat output of the system and its differential parameters. Replace the state variables with the transformation relationship of Equation (16) into the flat output form to obtain (X a , Ya , the multi-objective optimization model of the (x, l) trajectory is Equation (46).

[0221] The present invention uses the INSGA-II algorithm to solve the multi-objective optimization problem. In addition, the Pareto optimal solution set of the multi-objective optimization provides multiple optimal solutions. To obtain solutions that meet different objective priority requirements, the normalized weighted objective function is set as Equation (47), with f op minimized as the objective to select a balanced optimal solution. Different weights λ1, λ2, λ3 can be set according to different task requirements or user preferences to adjust the optimal solution.

[0222] 4 Simulation Experiments

[0223] To verify the effectiveness of the algorithm proposed in this paper, numerical simulation experiments are carried out on the MATLAB software based on the complete dynamic model. In the experiment, a custom two-dimensional 45*45 obstacle map is used. This map sets various obstacle models, including circles, triangles, and rectangles, and diverse layouts are considered in the design. Among them, the map not only contains relatively open areas but also focuses on simulating relatively crowded and narrow channels to reflect the complex environments that may be encountered in practical applications. The physical parameter values of the tower crane system are m = 1 kg, M = 5 kg, g = 9.8 m / s 2 , J = 6.8 kg·m 2 , and the initial and final positions are (x0 = 10 m, l0 = 10 m, θ 10 = 15°), (x d = 48 m, l d = 25 m, θ 1d = 61°).

[0224] 4.1 Tower Crane Path Planning Experiment

[0225] The parameters of the AD-BI-RRT algorithm are set as: O = 10 cm, X lim = 50 cm, α1 = 2, α2 = 10, d0 = 5 cm, m e = 2, the maximum step size limit is 3.5 cm, the minimum step size limit is 0.2 cm, the maximum number of iterations is 2000, and the hyperparameters of the DDPG algorithm are shown in Table 1.

[0226] Table 1

[0227] Parameter Value Actor network learning rate 0.003 Critic network learning rate 0.003 Experience replay buffer capacity 500000 Random experience mini - batch capacity 128 Initial value of noise standard deviation 0.1 Decay rate of noise standard deviation 0.995 Minimum noise standard deviation 0.01

[0228] The reward curve after training the DDPG algorithm 1500 times is as Figure 7 shown.

[0229] Three comparison algorithms, namely RRT*, improved BI-RRT*, and APF-RRT*, were selected to evaluate the performance of the AD-BI-RRT algorithm. The maximum number of iterations in the comparison algorithms was 2000, the fixed step size was 1.75 cm, the maximum connection distance was 5 cm, and the reconnection radius optimized by RRT* was 5 cm.

[0230] The simulation results of the four algorithms are compared as Figures 8 - 11 shown. Twenty experiments were conducted and the running time, path length, and number of nodes of the algorithms were recorded respectively. The test results are as Figure 8 shown in Figs. a-8d, and the experimental statistical results are as Figures 9 - 11 .

[0231] From Figure 8 the comparison of the four subgraphs, it can be seen that the AD-BI-RRT algorithm has good self-adaptability and can automatically adjust parameters in crowded or narrow areas to find a passable path. From Figure 9 it can be seen that the proposed algorithm is significantly less than the comparison algorithms in terms of running time. From Figure 10 and Figure 11 it can be seen that in terms of path length and number of nodes, AD-BI-RRT is significantly less than the comparison algorithms. It is improved by 78.8%, 21.7%, and 66.7% respectively in terms of running time, 23.7%, 16.2%, and 14.7% respectively in terms of path length, and 78.6%, 33.8%, and 66.9% respectively in terms of the average number of nodes. The efficient characteristics of the proposed AD-BI-RRT algorithm in the search process are verified.

[0232] 5.2 Selection of key path points

[0233] In the simulation experiment, the characteristic nodes after removing redundant nodes are (9.66, 2.59), (16.73, 11.16), (14.82, 24.98), (24.97, 33.72), (25.67, 39.77), and (23.21, 41.98) respectively. According to Equation (24), all path points are calculated and used as the profile points for subsequent trajectory planning. The specific values are shown in Table 2.

[0234] Table 2 Load path points

[0235]

[0236] 5.3 Tower crane trajectory planning experiment

[0237] 5.3.1 Multi-objective optimization based on NURBS curve

[0238] Taking the 6 path points obtained from path planning (shown in Table 3) as the profile points for the tower crane system trajectory planning. The corresponding multi-objective optimization decision variable X is expressed as Equation (43) according to Equation (34).

[0239] The trajectory optimization adopts the INSGA-II algorithm, and its time complexity is O(m·p·q), where m is the population size, p is the number of iterations, and q is the dimension of the optimization variables. In the experiment, m = 50 and p = 50 are set, and the trajectory optimization takes 35 seconds (MATLAB platform). Combining with the tower crane control period (100 - 500 milliseconds), real-time control can be achieved through the combination of offline planning and online tracking. In the future, using FPGA or GPU acceleration can further compress it to the millisecond level.

[0240] In the experiment, the optimal solution weights are set as λ1 = 0.8, λ2 = 0.1, and λ3 = 0.1. The state constraints are set as v 1lim = 2.5 m / s, v 2lim = 3 m / s, v 3lim = 7° / s, θ lim = 4°, a 1lim = 0.3 m / s 2 , a 2lim = 2 m / s 2 , a 3lim = 5° / s 2 . The obtained optimal solution vector is as shown in Equation (44).

[0241] The trajectory curve corresponding to the optimal solution vector is used as the target trajectory of the tower crane system. To verify the effectiveness of the optimization result, a traditional PID controller is used for trajectory tracking, and the obstacle avoidance trajectory of the load of the tower crane system is as Figure 12 shown.

[0242] It can be seen from Figure 12 that during the hoisting process, the load can accurately track the planned trajectory and does not collide with any obstacles.

[0243] 5.3.2 Comparative Experiment - B-Spline Curve

[0244] To evaluate the performance of the trajectory, a comparative experiment is set up. Under the same time interval, the B-spline curve method is used for trajectory planning, and the experimental results of the tower crane tracking are as Figures 13 to 15 .

[0245] It can be seen from Figures 13 to 15 that the curves all converge to the target position at about 5 s, and the speed and acceleration of the system all meet the set constraints. The response curves of the load swing angles θ2 and θ3 of the comparative algorithm still have obvious residual swings after reaching the target position. In contrast, the trajectory generated by the trajectory planning method proposed in this paper makes the load swing angle response significantly smaller, and the residual swing is close to 0.

[0246] The performance indicators are calculated. The total hoisting time, energy consumption, and swing angle evaluation function values of the comparison algorithm are 5.24 s, 1.58 J, and 0.0415 rad respectively. The hoisting time, energy consumption, and swing angle evaluation function values of the method proposed in the present invention are 5.24 s, 1.47 J, and 0.0313 rad respectively. It can be seen that under the condition of the same total hoisting time, the energy consumption of the algorithm proposed in the present invention is reduced by about 7%, and the anti-swing performance is improved by about 24.6%, having better performance.

[0247] 5.3.3 Comparative Experiment - EI Input Shaping

[0248] To further verify the superiority of the proposed algorithm, it is necessary to compare the algorithm proposed in this paper with other open-loop control methods. Therefore, the EI input shaping control (Extra Insensitive Input Shaper) of other tower cranes is used as a comparative experiment. In the experiment, the same trajectory as the present algorithm is selected for the rope length, and the experimental results of the tower crane tracking are as follows Figures 16 to 18 .

[0249] The performance indicators are calculated. The hoisting time, energy consumption, and swing angle evaluation function values of the EI shaping algorithm are 5.84 s, 1.52 J, and 0.0437 rad respectively. The algorithm in this paper reduces the hoisting time by 10.3%, reduces the energy consumption by 3.3%, and improves the anti-swing performance by 28.4%. Therefore, compared with the EI shaping control, the algorithm proposed in this paper has obvious advantages in control performance.

[0250] The above are only the preferred embodiments of the present invention, and do not impose any formal limitations on the present invention. Any simple modifications, equivalent changes, and decorations made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A motion planning method for tower cranes in complex environments, characterized in that: The method includes the following steps: (1) Based on differential flatness analysis, transform the motion planning problem of the system in the complex obstacle space into a planning problem of the flat output; (2) Generate a reference path based on the Directional Biased Bidirectional Rapidly-exploring Random Trees (DB-BiRRT*) algorithm with direction bias, and introduce a direction guidance mechanism to improve the efficiency and quality of path planning; (3) Fit the hoisting trajectory based on the reparameterization of the Non-Uniform Rational B-Spline (NURBS) curve, optimize the fitted trajectory, and considering the full-state constraints of the system such as load swing reduction and obstacle avoidance, pose a multi-objective optimization problem with the goals of minimizing the hoisting time, energy consumption, and swing. Optimize the shape parameters and time node distribution of the NURBS curve to obtain the comprehensive optimal trajectory curve of the system state, and use the Non-dominated Sorting Genetic Algorithm II (NSGA-II) for optimization and solution; The step (2) includes the following sub-steps: (2.1) Based on the Directional Biased Bidirectional Rapidly-exploring Random Trees (BiRRT) algorithm, guide the node generation process by introducing a target bias mechanism that fuses regional probability sampling and a direction guidance mechanism that improves the potential field function, so that new nodes are generated in a better direction; (2.2) Construct a parameter decision link based on the Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm, design a reward function, and dynamically generate important parameters in the proposed random tree node expansion process, thereby constructing an adaptive bidirectional RRT algorithm with direction bias; (2.3) When selecting path points of the tower crane system load, use the greedy algorithm to eliminate redundant points and only retain key path points.

2. The tower crane motion planning method in a complex environment according to claim 1, characterized in that: In the said step (1), according to formula (16), all state variables of the system are represented by (X a , Y a , x, l) and their differential variables. The system is differentially flat. For the motion planning of the tower crane system, it is equivalently transformed into the motion planning of the flat output (X a , Y a , x, l). Among them, are the conversion relationships between the states θ1, θ2, θ3 and the flat output.

3. The tower crane motion planning method in a complex environment according to claim 2, characterized in that: In the step (2.1), when expanding the random tree, there is a probability of directly expanding towards the target point, thereby accelerating the expansion speed towards the target point: Among them, q rand , q goal , q pro are random sampling points, target points, and regional probability sampling points, P0 ∈ (0, 1) is the bias probability, and P ∈ (0, 1) is a random number; In the case where the target bias is not executed when P > P0, when generating sampling nodes, divide the entire map into countless random sampling regions, and based on the line connecting the starting point and the target point, calculate the sampling probability of each region using the distance from the random sampling region to this line as a variable, and use the Gaussian distribution to determine the sampling probability in each region where d c is the distance from the center point of the region to the line connecting the start and end points, l is the distance between the start and end points, and the standard deviation of the Gaussian distribution probability P(d c ) is 0.25l, p st , p end represent the start point and the end point respectively, and p c is the center point of these sampling regions.

4. The tower crane motion planning method under complex environments according to claim 3, wherein: In the step (2.1), the improved potential field function is: Among them, α i (i = 1, 2), β are the attraction coefficient and the repulsion coefficient respectively, q, q o , q g , q r represent the current point, the obstacle, the target point and the sampling point respectively, d(q o ), d g (q g ), d r (q r ) are the Euclidean distances between the current point and the obstacle, the target point and the sampling point respectively, d a is the limiting range of the maximum target attraction, d0 is the maximum distance at which the potential field takes effect, m is the repulsion intensity parameter, and m e > 0; The attraction and repulsion functions are the negative gradients of the potential field function, and the attraction force F of the current node att and the repulsion force F req : F att2 = α2d r (q, q r ) F att = F att1 + F att2 (21) Among them, F att1 The direction is from the current node to the target point, F att2 The direction is from the current node to the sampling point, F req The direction is from the obstacle to the current node.

5. The tower crane motion planning method in a complex environment according to claim 4, characterized in that: In the step (2.2), model the constructed random tree expansion mechanism as the environment of reinforcement learning, including the following: (2.2.1) State space: In order to obtain sufficient environmental information of the nodes, the state space is designed as: S = {D g1 , D g2 , D r1 , D r2 , D ob1 , D ob2 , N p1 , N p2} (23) Among them, D g1 , D g2 are respectively the 2D vector values from the current nodes of the two trees to the corresponding target points, D r1 , D r2 are respectively the 2D vector values from the current nodes of the two trees to their nearest randomly sampled points, D ob1 , D ob2 is the 2D vector value from the current nodes of the two trees to their nearest obstacle, N p1 , N p2 is the total number of nodes within the range O of the current nodes of the two trees, and the entire state space has a total of 14 dimensions; (2.2.2) Action space: According to the design and parameter requirements of the improved random tree expansion mechanism, it is designed as: A = {P1, P2, β1, β2, L1, L2} (24) where P1 and P2 are the bias probabilities of the two trees towards the target point respectively, β1 and β2 are the current repulsive force coefficients of the two trees, and L1 and L2 are the current expansion steps of the two trees. The entire action space has a total of 6 dimensions; (2.2.3) Reward function: Considering the sparse reward problem and the optimization effect comprehensively, the reward function design includes: obstacle avoidance reward when the random tree expands nodes, proximity reward when the random tree expands nodes, connection reward after obtaining a feasible path, and path optimization reward after obtaining a feasible path.

6. The tower crane motion planning method under complex environments according to claim 5, wherein: In the step (2.2.3), Obstacle avoidance reward: Each time an extended node collides with an obstacle or is less than the safe distance, a reward value of -1 is given, and the reward is 0 at other times, which is expressed as: where d ob is the distance to the nearest obstacle, and d safe is the set safety distance; Approaching reward: When the generated node is far from the target point, a reward value of -1 is given, and the reward is 0 at other times, which is expressed as: where d goal is the distance from the current node to the target point, and d b is the distance from its parent node to the target point; Connection reward: If two random trees are successfully connected to generate a feasible path, a reward of 100 is given, and if the connection fails, a reward value of -200 is given, which is expressed as: Path optimization reward: Reward is given based on the length of the obtained path. If the path exceeds the limit value X, lim a penalty is imposed. When the path is the current global optimum and does not exceed the limit, a reward is given; otherwise, the reward is 0. It can be expressed as:

7. The tower crane motion planning method under complex environments according to claim 6, characterized in that: In the step (2.2), the training process of the decision-making model is as follows: (2.2.a) Initialize the Actor-Critic network, set the initial weights of the network and initialize the experience replay pool, and reset the reinforcement learning environment, including the state space, action space, reward, and random tree nodes; (2.2.b) The agent selects an action according to the current state and policy, obtains the next node according to the improved random tree expansion mechanism, and calculates the reward for this action according to (25)-(28). During the expansion process, record the current state, the action taken, the obtained reward, and the next state, and store them in the experience replay pool; (2.2.c) Randomly extract a certain number of samples from the experience replay pool, and update the weights of the actor and critic networks by the gradient descent method. In each update, the goal is to maximize the Q value output by the Critic network, thereby optimizing the policy selection of the agent; (2.2.d) Repeat the above training steps. As the training progresses, the policy of the agent will be gradually improved and tend to be optimized. Finally, the reward curve tends to be stable and the training is completed.

8. The tower crane motion planning method in a complex environment according to claim 7, characterized in that: In the step (3), the fitted (X a , Y a , x, l) trajectory is optimized, and the multi-objective optimization decision variable X is designed as follows: In the formula, ω is the weight factor vector of the curve, t is the value point time interval vector, and the number of elements in the two vectors are 2z + 14 and z - 1 respectively; During the hoisting process, the state constraints of the operating room: Constrain the swing angle and its speed: In the formula, v ilim (i = 1,..., 5) is the maximum value constraint of the velocity of each physical quantity, a ilim (i = 1, 2, 3) is the maximum value constraint of the acceleration of the all-driven physical quantity, θ lim is the maximum value constraint of the swing angle; Let Ω be the set of non-obstacle positions, and constrain the generalized coordinates of the load: (x, l, θ1, θ2, θ3) ∈ Ω (42) Select three objective functions for the multi-objective optimization problem: In the formula, f0(x) is the hoisting time, v x (t), v θ1 (t) and v l (t) are the luffing displacement, slewing angle and rope length speed respectively, f1(x) is the total energy consumption of the actuator, f2(x) is the evaluation function of the swing angle, x, y ∈ (0, 1) are the user-defined weights, and max|θ j |(j = 2, 3) represents the maximum absolute value of the swing angle.

9. The tower crane motion planning method in a complex environment according to claim 8, characterized in that: In the said step (3), the state variables are replaced with a flat output form formula according to the conversion relationship of formula (16), and the multi-objective optimization model of the trajectory of (X a , Y a , x, l) is as follows: Among them, g1~g5 are the constraints satisfied by the optimization problem; To obtain solutions that meet different objective priority requirements, the normalized weight objective function is set as formula (47), with f op minimized as the objective to select a balanced optimal solution, and different weights λ1, λ2, λ3 are set according to different task requirements or user preferences to adjust the optimal solution In the formula, n is the number of individuals in the population, and f 0i , f 1i , f 2i is the objective function value of the i-th individual of the above three objective functions. f 0min , f 1min , f 2min is the minimum value of each of the three objective functions among all individuals. f 0max , f 1max , f 2max is the maximum value of each of the three objective functions among all individuals. λ1, λ2, and λ3 are the weights of the three objective functions, which are set according to preferences or task requirements as needed.

10. Tower crane motion planning device under complex environment, characterized in that: It includes: A differential flatness analysis module, which based on differential flatness analysis, transforms the motion planning problem of the system in a complex obstacle space into a planning problem of flat output; A path planning module, which generates a benchmark reference path based on the direction-biased BiRRT* algorithm DB-BiRRT, and introduces a direction guidance mechanism to improve the efficiency and quality of path planning; A trajectory planning module, which performs hoisting trajectory fitting based on non-uniform rational B-spline NURBS curve reparameterization, optimizes the fitted trajectory, and under the full state constraints of the system considering load swing reduction and obstacle avoidance, with the goal of minimizing hoisting time, energy consumption, and swing, proposes a multi-objective optimization problem to optimize the shape parameters and time node distribution of the NURBS curve, and obtains the comprehensive optimal trajectory curve of the system state, and uses the non-dominated sorting genetic algorithm NSGA-II for optimization and solution; The path planning module executes: (2.1) BiRRT algorithm based on direction bias guides the node generation process by introducing a target bias mechanism that fuses regional probability sampling and a direction guidance mechanism that improves the potential field function, enabling new nodes to be generated in a more optimal direction. (2.2) Construct a parameter decision-making link based on the DDPG reinforcement learning algorithm to dynamically generate important parameters in the proposed random tree node expansion process, thereby constructing an adaptive bidirectional RRT algorithm based on direction bias and designing a reward function. (2.3) When selecting path points for the tower crane system load, the greedy algorithm is used to eliminate redundant points and only key path points are retained.

Citation Information

Cited By

  • Bridge crane path planning method considering obstacle avoidance

    CN120736412A

  • Deep reinforcement learning driven crane trajectory planning method

    CN122221699A