Intelligent tower crane path planning method based on reinforcement learning and imitation learning

By combining reinforcement learning and imitation learning techniques, an intelligent path planning method for tower cranes was developed, which solved the problems of safety risks and path planning efficiency in tower crane operations, and achieved the generation of safe, effective and efficient lifting paths.

CN120069250APending Publication Date: 2025-05-30CHINA CONSTR FIRST DIV GROUP CONSTR & DEV +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411931133.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Tower crane operation poses a great safety risk, and existing path planning algorithms are difficult to generate safe, effective and efficient lifting paths in complex environments, especially in dynamic environments and high-dimensional spaces.

Method used

Combining reinforcement learning and imitation learning technology, a tower crane path planning method was developed. Through behavioral cloning algorithms, near-end strategy optimization algorithms and generative adversarial imitation learning algorithms, an intelligent tower crane path planning model is built to generate a lifting path that has both safety, operability and efficiency.

Benefits of technology

It realizes the generation of safe, effective and efficient tower crane lifting paths in complex environments, improves the intelligence and automation level of tower cranes, and reduces the impact of human factors on operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069250A_ABST
    Figure CN120069250A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent tower crane path planning method based on reinforcement learning and imitation learning. The method comprises the following steps: firstly, simulating a hoisting track demonstrated by an expert based on a behavioral cloning algorithm BC to learn a feasible initial hoisting strategy; and then carrying out linear combination on loss functions of a near-end strategy optimization algorithm PPO and a generative adversarial imitation learning algorithm GAIL based on a training strategy of a linear function so as to fully combine the advantages of reinforcement learning and imitation learning and finally construct an intelligent path planning model of the tower crane, so that the operability of path planning is improved, and the path planning efficiency is improved. The invention provides a path planning model combining a reinforcement learning algorithm and an imitation learning algorithm, through learning and imitation of a hoisting strategy of an operator, interaction with the environment and continuous iterative optimization are carried out in training, and a tower crane hoisting path with relatively high transportation efficiency, hoisting safety and operability at the same time is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent path planning method for tower cranes based on reinforcement learning and imitation learning, belonging to the technical fields of engineering construction management and artificial intelligence optimization. Background Art

[0002] The tower crane, also known as the tower hoist or tower crane, is an indispensable mechanical equipment on modern construction sites, especially playing a crucial role in the construction of prefabricated buildings. During the prefabricated construction process, the tower crane can perform material transportation tasks in the vertical and horizontal directions within a large coverage area, thus playing a key role in improving the transportation efficiency of the construction site and ensuring the project progress. [1] However, the tower crane lifting operation has characteristics such as large weight of the lifted materials, high lifting height, and overlapping working areas, which makes the tower crane operation have relatively high safety risks. Usually, the occurrence of tower crane safety accidents will cause serious casualties and property losses. In 2022, a total of 108 special equipment accidents occurred nationwide, resulting in 101 deaths. [2] Classified by equipment type, there were 25 lifting machinery accidents, accounting for 23.15% of the total accident types. Classified by occurrence link, 95 accidents occurred in the use link, accounting for 87.96% of all occurrence links. Among them, 80% of the tower crane accidents in construction projects are caused by human factors such as operation errors, improper command, and illegal operations, fully exposing the limitations of the current tower crane operation relying on manual experience and manual control. At the same time, human factors such as fatigue driving and lack of operation experience have a greater impact on the lifting path when encountering sudden or complex situations. In order to improve construction efficiency and construction safety, tower cranes are gradually developing towards the direction of intelligence and automation, including automatic tower crane model selection, positioning, and lifting path planning.

[0003] In the field of tower crane automatic control research, path planning has always been a challenging research. The concept of path planning was first proposed in the 1960s and is mainly applied to the problem of robot obstacle avoidance behavior planning. [3] The accuracy and effectiveness of path planning mainly depend on the description method of space and objects in the optimization model. [4] Related research generally regards the tower crane as a three-dimensional kinematic model with multiple degrees of freedom. [5,6], since in path planning, it is necessary to consider both the collision between the hoisted materials and the surrounding environment and the collision between the components of the crane and the surrounding environment, the complexity of describing the tower crane's attitude in the Cartesian coordinate system is relatively high. Therefore, the concept of configuration space (C-Space) in robot planning theory is introduced. By representing each coordinate axis as a degree of freedom of the tower crane, the complex attitude description can be transformed into the coordinates of a particle in the configuration space, thereby reducing the complexity of the problem. To prevent collisions between objects, related research has designed hierarchical bounding volumes, such as axis-aligned bounding box [7] (Axis-alignedBounding Box,AABB), oriented bounding box [8] (Oriented Bounding Box,OBB) and sphere bounding box [9] (Sphere Bounding Box,SBB) and other methods, which wrap objects with complex geometric shapes with a regular geometric shape, thereby accelerating the collision detection operation. Thus, the problem of optimizing the hoisting path of the crane can be transformed into: searching for the shortest path from the material stacking position to the unloading position on the premise of avoiding the overlap of each object along each axis in the configuration space. Traditional deterministic algorithms, such as Dijkstra

[10] and A*

[11] , can find the optimal hoisting path by iteratively selecting the node with the currently shortest path estimate

[12] . However, these algorithms need to search the entire configuration space to find the optimal solution. As the space dimension increases, the computational amount will also increase, making it difficult to apply to complex environments and high-dimensional space problems.

[0004] Some scholars have adopted sampling-based path planning algorithms, such as probabilistic roadmap

[13] (PRM), rapidly-exploring random tree [14-18] (RRT), rapidly-exploring random tree star [19 (RRT*) and bidirectional rapidly-exploring random tree [20-22] (Bi-RRT) algorithms to improve the path planning speed. PRM first generates a set of obstacle-free sampling points through a random sampling strategy, then connects two sampling points that are at a connected distance to establish an undirected roadmap, and finally searches for the shortest path connecting the starting point and the target point on this graph through a graph search algorithm

[23] . The calculation time of PRM depends on the number of path nodes, which makes it more suitable for finding feasible paths rather than optimal paths in complex and high-dimensional spaces

[13] . Compared with PRM, the advantage of RRT is that it is a single planning algorithm

[24] This means that the RRT is designed to run separately for each planning task, generating a path from the starting point to the ending point without the need to know the global information of the environment in advance, thus significantly shortening the search time. However, the paths generated by the RRT-based path planning algorithm have defects such as many inflection points and low path smoothness. Too many inflection points will increase the swinging amplitude of the material in the air, making it much more difficult for the tower crane operator to execute the lifting path.

[25] To improve the path quality and search efficiency, Zhou et al.

[16] improved the original RRT algorithm using the method of generalized distance and cells, and the results showed that the quality of the planned path was improved when the planning time was similar to that of the RRT algorithm. Lin et al.

[20] used a bidirectional rapidly-exploring random tree to perform path planning for a tower crane model with 7 degrees of freedom. Hu et al.

[18] proposed an improved RRT* strategy, which can improve the practicality and safety of the planned path. The above algorithms rely to a large extent on parameter optimization in the mathematical model and lack the autonomy and intelligence to complete tasks. Facing complex tasks in a dynamic environment, reliable and efficient artificial intelligence algorithms that can learn from experience, adapt to changing conditions, and make autonomous decisions need to be developed. In recent years, with the rapid development of artificial intelligence algorithms such as Reinforcement Learning (RL), there is still a large research space for the tower crane path planning problem. Reinforcement learning continuously iteratively learns by interacting with the environment and receives reward feedback, updating and adjusting the strategy to obtain the maximum reward value. This process does not require prior knowledge and precise modeling of the environment and has good adaptability to dynamic environments. However, when facing practical problems, it is difficult for reinforcement learning to define a suitable reward function, and usually, it is difficult to achieve the expected behavior only by designing task rewards. The tower crane path planning problem is a multi-objective optimization problem that needs to consider the safety, practicality, and planning efficiency of the planned path simultaneously.

[0005] In summary, aiming at the actual needs of engineering projects, this study focuses on proposing a tower crane path planning method combining imitation learning and reinforcement learning to generate a collision-free lifting path that takes into account both operation feasibility and efficiency, which has a certain promoting effect on improving the intelligence of tower cranes. Summary of the Invention

[0006] The essence of the path planning problem is to find a collision-free path from the starting point to the target point. However, tower crane path planning is a multi-objective optimization problem. When considering safety, the planned path also needs to balance and optimize the objectives of efficiency, energy consumption, and practicality. The potential conflicts between these objectives require the path planning algorithm to not seriously damage the achievement of other objectives when achieving one objective. Based on the basic theories of reinforcement learning and imitation learning, this study combines the two to establish a path planning model, which aims to generate a tower crane lifting path that is safe, operable, and has high lifting efficiency. The research will be carried out in the following order:

[0007] First, the model is pre-trained based on the behavior cloning algorithm. After obtaining the initial policy, the proximal policy optimization algorithm and the generative adversarial imitation learning algorithm are used to interact with the environment to further optimize the policy. Finally, a tower crane intelligent path planning model is constructed. For an actual engineering case, the established tower crane intelligent path planning model is tested. The tower crane path planning model includes two stages: pre-training and model fine-tuning. The modeling process is as follows:

[0008] 1) Construction of a virtual simulation environment for tower crane operation;

[0009] To develop the tower crane driving simulation function using VR devices and accessories, basic environment configuration for the HTC Vive Pro 2 head-mounted display and the T.16000M joystick needs to be carried out through Unity3D to achieve connection with Unity and build a tower crane driving simulator. The specific steps are as follows:

[0010] Step 1: Through research on the actual working environment of tower crane drivers, understand the size of the tower crane cab space, the layout of the joystick, and the functions of the buttons, etc., and plan the operation area of the tower crane driver.

[0011] Step 2: Install the left and right positioning base stations of the HTC Vive Pro to ensure accurate tracking of the head-mounted display.

[0012] Step 3: In the Unity top toolbar, select Window -> Asset store to open the Unity Asset store, download the Steam VR plugin from it and install it into the Unity project.

[0013] Step 4: Add the "Camera Rig" component of Steam VR at the position of the tower crane driver's perspective height in the virtual environment tower crane cab, set the relevant parameters, and define the perspective of the tower crane driver in the virtual environment, including the position and orientation of the head.

[0014] Step 5: Select Edit -> Project Settings -> Input Manager in the top toolbar of Unity, and add three new axes for the joystick, named "HookVertical", "HookHorizontal", and "Rotation" respectively to correspond to the movements of the tower crane in three degrees of freedom. Then configure the parameters of these axes to match the corresponding inputs of the T.16000M joystick. Set the "Type" to "Joystick Axis", and map the axes of the T.16000M joystick to the corresponding axes in the Input Manager according to the actual functions of the tower crane joystick (the x-axis of the left joystick controls the slewing of the tower crane, the y-axis controls the luffing of the trolley, and the y-axis of the right joystick controls the lifting and lowering of the hook).

[0015] Step 6: Connect the power supplies of accessories such as the base station and the streaming box to make the head-mounted display in a normal working state.

[0016] Step 7: Connect the T.16000M joystick and change the button control of the tower crane dynamic model to axis control.

[0017] Step 8: Click to run the program in Unity 3D. The operator performs an immersive tower crane simulation hoisting task through the head-mounted display.

[0018] 2) Design of the state space of the tower crane agent;

[0019] The state space is represented as the set of all possible states that can be observed by the environment at each time step. During the decision-making process, the tower crane agent receives the state information at a certain moment and then selects an action from the action space to control the movement of the tower crane.

[0020] The state information of the tower crane itself consists of the parameters included in formula (1):

[0021] S crane =[θ t ,h t ,d t ,r t ,c t ,g t ,b t ,l t (1)

[0022] where θ t represents the rotation angle of the boom at time step t; h t represents the height of the tower crane hook from the ground at time step t; d t represents the radial distance between the slewing center of the tower crane and the luffing trolley at time step t; r t represents the ratio of the current moving distance to the distance between the starting point and the target point at time step t; ct represents the direction vector from the hook position to the target point position at time step t; g t represents the dot product value between the direction vector and the ground normal vector; b t is a binary value used to determine whether the lifting task is completed at time step t; l t represents the position information of the material in three-dimensional space at a certain moment. The specific description of the state information is shown in Table 1.

[0023] Table 1 Tower crane self-state information

[0024]

[0025] 3) Design of the tower crane agent action space;

[0026] The action space can be divided into two categories: continuous action space and discrete action space. Since the tower crane path planning problem is essentially a continuous space problem, according to the analysis of the kinematic characteristics of the tower crane, the movement of the tower crane with three degrees of freedom is transformed into continuous action values, and the action space is expressed as:

[0027] A = [T t , J t , H t (2)

[0028] T t , H t and J t are the action vectors of the tower crane movement, reflecting the kinematic characteristics of the tower crane, that is, the luffing movement of the trolley, the hoisting movement of the hook, and the slewing movement of the boom. Normalize the actions of the tower crane and limit them between [-1, 1] after selecting the actions. Formula (3) represents the specific normalization method:

[0029]

[0030] where A norm ∈[-1, 1], A ∈ [A min , A max , A min and A max represent the minimum and maximum values of the corresponding elements in the action variables respectively.

[0031] 4) Design of the tower crane agent reward function;

[0032] For the tower crane path planning problem, the reward function is designed from the following aspects: First, encourage the tower crane agent to approach the target position, that is, continuously reduce the distance between the tower crane hook and the target point. Therefore, the change in the distance between the hook at the current and the previous time steps and the target position is used as part of the reward function, expressed in the form of formula (4):

[0033] R 1 = clip(k distance (P t-1 - P t ), - 0.1, + 0.1) (4)

[0034] Among them, P t-1 represents the distance between the tower crane hook and the target position at step t - 1, and P t represents the distance between the tower crane hook and the target position at time step t. k distance is a constant coefficient. The distance reward value at each time step is limited between - 0.1 and 0.1 through the clip function. As shown in formula (5), the calculation formula for the distance between the hook and the target position is:

[0035]

[0036] Among them, (x hook , y hook , z hook ) is the current position of the hook in three - dimensional space, and (x goal , y goal , z goal ) is the target point of the tower crane lifting task.

[0037] Similarly, it is necessary to additionally encourage the tower crane boom to rotate towards the target point position, that is, continuously reduce the angle between the boom and the target point. Therefore, formula (6) takes the angle change between the current and the previous time step of the boom and the target position as the boom orientation reward function, which can be expressed as:

[0038] R 2 = clip(k direction (D t-1 - D t ), - 0.1, + 0.1) (6)

[0039] Among them, D t-1 represents the angle between the boom and the target point in the x - y plane at time step t - 1, and D t represents the angle between the boom and the target point in the x - y plane at time step t. k direction is a constant coefficient. The distance reward value at each time step is limited between - 0.1 and 0.1 through the clip function.

[0040] The obstacles at the construction site are mainly of two types: buildings and materials stacked at the construction site. The reward function also includes the collision judgment between the tower crane body and these two types of obstacles. In actual tower crane lifting operations, for safety reasons, it is often required to maintain a certain safety distance between the tower crane and the obstacles. Considering the complexity of the virtual environment and the operating speed of the tower crane, the radius of the safety buffer is set to 5m. That is, when the distance between the tower crane and the obstacle is less than 5m, it is considered a collision, and at this time, the agent receives a discrete reward R of -2. 3 。

[0041] Secondly, the reward function also includes the judgment of the tower crane lifting materials reaching the target point. When the lifted materials reach the target position, a discrete reward R of +10 is received. 4 Among them, r is the task success radius of the target point, which is set to 2m. That is, when the distance between the lifted materials and the target point is less than r, it is considered that the lifted materials have successfully reached the target point, and the lifting task is completed.

[0042] Finally, in order to encourage the tower crane agent to complete the path planning task as soon as possible, that is, to transport the lifted materials to the target point using as few time steps as possible, a very small step penalty is designed. The tower crane agent will receive a time step penalty shown in formula (7) for each step during the training process, where T maxstep represents the maximum number of steps for the tower crane to perform the lifting task, which is expressed as:

[0043]

[0044] 5) Algorithm process design;

[0045] In the BC+PPO+GAIL model, the behavior cloning algorithm is used to train the initial policy network by minimizing the difference between the actions generated by the policy and the expert actions. The network parameters are updated using the method of gradient descent. To minimize the loss function θ of behavior cloning. The parameter update formula is expressed as (8):

[0046]

[0047] In the model fine-tuning stage, the network parameters of the behavior cloning strategy are used as the initial parameters of the PPO and GAIL policy networks. A weight coefficient λ(N) related to the number of iterations N is defined to balance the optimization of the two algorithms for the policy network. The formula (9) represents the expression of the hybrid loss function:

[0048] L mix =λ(N)L GAIL (θ)+(1 - λ(N))L PPO (φ) (9)

[0049] Among them, λ(N) is the weight coefficient, which is used to balance the losses of PPO and GAIL. L GAIL L(θ) represents the loss function of GAIL, which focuses on the effect of imitating the expert policy; L PPO L(φ) represents the loss function of PPO, which focuses on policy optimization through a pre-set reward function.

[0050] During the training process, λ(N) should continuously decrease as the number of iterations increases. The PPO algorithm helps to maintain the stability of the learned behavior, reduce the possibility of catastrophic forgetting of the agent during further exploration, and at the same time further improve the generalization ability of the model by continuously interacting with the environment. Therefore, a linear decay strategy is used to reduce the weight coefficient. First, set an initial weight λ start and the final desired weight λ goal . Then, calculate the size of the weight decay for each iteration according to the total number of iterations N of the model training. The amount of weight decayed each time is shown in formula (10):

[0051]

[0052] In each iteration, update the weight coefficient according to the amount of weight decayed Δλ calculated by formula (11):

[0053] λ' = max(λ - Δλ, λ goal ) (11)

[0054] where λ' represents the new weight coefficient and λ represents the old weight coefficient. The max function is used to ensure that λ will not be lower than the target value λ goal . In terms of parameter setting, set λ start to 0.8 and λ goal to 0.2 to ensure that the training can transition from imitation learning to trial-and-error learning of interacting with the environment. Description of the Drawings

[0055] Figure 1 Each component of the virtual construction site scene.

[0056] Figure 2 Development of the tower crane hoisting simulation function based on VR.

[0057] Figure 3 Flowchart of the implementation of this method.

[0058] Figure 4 Project construction layout plan.

[0059] Figure 5 Comparison of the actual path and the corresponding BC+PPO+GAIL planned path lengths.

[0060] Figure 6Visual comparison of the planned path and the actual path of BC+PPO+GAIL. (a) is the top view; (b) is the side view; (c) is the 3D view; Detailed implementation mode

[0061] The present invention will be described below in conjunction with the accompanying drawings and examples.

[0062] The model application and effect verification are as follows:

[0063] Taking the tower crane lifting task of an actual engineering project as an example, by comparing the results of the path planning model with the actual lifting path, the effectiveness of the proposed path planning model is verified. This case comes from an actual construction project, where two tower cranes are deployed inside the project site for the construction of two 7- to 9-story buildings. As Figure 4 shown, these two buildings are located in the core area of the construction site, and the safety monitoring system of the tower crane records the operation data of the tower crane in real time during the entire construction stage of the project.

[0064] The present invention uses the trained path planning model to execute 27 lifting tasks in the simulation environment. The starting and target points of these tasks correspond one by one to the starting and ending points of the 27 actual lifting paths collected. Table 2 shows the comparison of the results of the paths planned by BC+PPO+GAIL and the actual lifting paths. It can be seen that the average error of the lifting time of the BC+PPO+GAIL algorithm is 41.17 seconds. This time difference may be attributed to the extra time formed by workers hooking up and unloading building materials, emphasizing the complexity of the lifting process at the actual construction site, thus affecting the overall lifting time. Figure 5 shows the comparison of the lengths of the 27 lifting paths generated by BC+PPO+GAIL and the actual path lengths. The results show that the path length planned by the BC+PPO+GAIL algorithm is slightly longer than the actual lifting path, with an average error of 7.49 m.

[0065] Table 2 Comparison of the paths planned by BC+PPO+GAIL and the actual lifting paths

[0066]

[0067] Figure 6 shows the visual comparison between the actual lifting path and the path planned by BC+PPO+GAIL, where the blue path is the actual lifting path and the green represents the lifting path generated by the BC+PPO+GAIL strategy. It can be seen that BC+PPO+GAIL successfully avoids simultaneous hook lifting, tower arm slewing, and trolley luffing movements in the blind area, specifically manifested as the path visibility of the planned path being very close to that of the actual path. At the same time, BC+PPO+GAIL maintains a safe distance from the building during the lifting process, with a success rate of 100%.

[0068] References:

[0069] [1] Liu Meng, Huang Chun, Wang Jingjing, Wang Wenqi, et al. Optimization of tower crane selection and layout based on mixed-integer linear programming [J]. Journal of Civil Engineering and Management, 2020, 37(02): 142-150.

[0070] [2] Special Equipment Safety Supervision Bureau of the State Administration for Market Regulation, https: / / www.samr.gov.cn / tzsbj / qktb / tb / art / 2023 / art_67857a52cc3b41bd9cff0560e6b9b503.html.

[0071] [3] Raja P, Pugazhenthi S. Optimal path planning of mobile robots: A review [J]. International journal of physical sciences, 2012, 7(9): 1314-1320.

[0072] [4] Fang, Y., & Cho, Y.K. (2017). Effectiveness analysis from a cognitive perspective for a real-time safety assistance system for mobile crane lifting operations. Journal of Construction Engineering and Management, 143(4), 05016025.

[0073] [5] Kang S C, Miranda E. Planning and visualization for automated robotic crane erection processes in construction [J]. Automation in Construction, 2006, 15(4): 398-414.

[0074] [6] Kang S C, Miranda E. Computational methods for coordinating multiple construction cranes[J]. Journal of Computing in Civil Engineering, 2008, 22(4): 252 - 263.

[0075] [7] Choset H, Lynch K M, Hutchinson S, et al. Principles of robot motion: Theory, algorithms, and implementations[M]. Cambridge: MIT Press, 2005.

[0076] [8] Dutta S, Cai Y, Huang L, et al. Automatic re - planning of lifting paths for robotized tower cranes in dynamic BIM environments[J]. Automation in Construction, 2020, 110: 102998.

[0077] [9] Chang J W, Wang W, Kim M S. Efficient collision detection using a dual OBB - sphere bounding volume hierarchy[J]. Computer - Aided Design, 2010, 42(1): 50 - 57.

[0078]

[10] Soltani A R, Tawfik H, Goulermas J Y, et al. Path planning in construction sites: performance evaluation of the Dijkstra, A*, and GA search algorithms[J]. Advanced engineering informatics, 2002, 16(4): 291 - 303.

[0079]

[11] Sivakumar P L, Varghese K, Babu N R. Automated path planning of cooperative crane lifts using heuristic search[J]. Journal of computing in civil engineering, 2003, 17(3): 197 - 207.

[0080]

[12] Lin X, Han Y, Guo H, et al. Lift path planning for tower cranes based on environmental point clouds[J]. Automation in Construction, 2023, 155: 105046.

[0081]

[13] Chang Y C, Hung W H, Kang S C. A fast path planning method for single and dual crane erections[J]. Automation in Construction, 2012, 22: 468 - 480.

[0082]

[14] Chen Z, Min L, Shao X, et al. Obstacle avoidance path planning of bridge crane based on improved RRT algorithm[J]. Journal of system simulation, 2021, 33(8): 1832 - 1838.

[0083]

[15] Zhang C, Hammad A. Improving lifting motion planning and re - planning of cranes with consideration for safety and efficiency[J]. Advanced Engineering Informatics, 2012, 26(2): 396 - 410.

[0084]

[16] Zhou Y, Zhang E, Guo H, et al. Lifting path planning of mobile cranes based on an improved RRT algorithm[J]. Advanced Engineering Informatics, 2021, 50: 101376.

[0085]

[17] Kang, S.C., & Miranda, E. (2009). Numerical methods to simulate and visualize detailed crane activities. Computer-Aided Civil and Infrastructure Engineering, 24(3), 169 - 185.

[0086]

[18] Hu S, Fang Y, Guo H. A practicality and safety-oriented approach for path planning in crane lifts[J]. Automation in Construction, 2021, 127: 103695.

[0087]

[19] Lin Y, Wang X, Wu D, et al. Lift path planning for telescopic crane based-on improved hRRT[J]. International Journal of Computer Theory and Engineering, 2013, 5(5): 816 - 819.

[0088]

[20] Lin Y, Wu D, Wang X, et al. Lift path planning for a nonholonomic crawler crane[J]. Automation in Construction, 2014, 44: 12 - 24.

Claims

1. A tower crane path intelligent planning method based on reinforcement learning and imitation learning. The method first learns a feasible initial lifting strategy based on the lifting trajectory demonstrated by the expert using the behavior cloning algorithm BC; Then, based on the linear function training strategy, the loss functions of the proximal policy optimization algorithm PPO and the generative adversarial imitation learning algorithm GAIL are linearly combined to fully combine the advantages of reinforcement learning and imitation learning, and finally a tower crane intelligent path planning model is constructed to improve the operability of the planned path. It is characterized in that: the tower crane path planning model includes two stages: modeling and model fine-tuning; the modeling process is as follows: 1) Construction of virtual simulation environment for tower crane operation; Use VR devices and accessories to develop tower crane driving simulation functions, configure the head-mounted display and joystick through Unity3D, connect with Unity, and build a tower crane driving simulator; 2) Design of tower crane agent state space; The state space is represented as the set of all possible states observed by the environment at each time step. In the decision-making process, the crane agent receives the state information at a certain moment, and then selects actions from the action space to control the crane movement. 3) Design of action space for crane agent; The action space is divided into two categories: continuous action space and discrete action space. Since the tower crane path planning problem is essentially a continuous space problem, the three-degree-of-freedom tower crane motion is converted into a continuous action value based on the kinematic characteristics analysis of the tower crane. 4) Design of reward function for crane agent; For the tower crane path planning problem, the reward function is designed from the following aspects: first, the tower crane agent is encouraged to get closer to the target position, that is, the distance between the tower crane hook and the target point is continuously reduced, and the change in the distance between the hook and the target position at the current and previous time steps is used as part of the reward function; The tower crane arm is additionally encouraged to rotate toward the target point, that is, the angle between the tower arm and the target point is continuously reduced, and the angle change between the tower arm and the target position at the current and previous time steps is used as the tower arm direction reward function; Secondly, the reward function also includes the judgment of whether the tower crane has reached the target point. When the crane reaches the target point, it receives a discrete reward R4 of +10. r is the mission success radius of the target point, which is set to 2m. That is, when the distance between the crane and the target point is less than r, it is considered that the crane has successfully reached the target point and the lifting mission is completed. Finally, in order to encourage the crane agent to complete the path planning task as quickly as possible, that is, to transport the hoisted materials to the target point in as few time steps as possible, a step penalty is designed; the crane agent will receive a time step penalty as shown in formula (7) for each step during the training process, where T maxstep It indicates the maximum number of steps that the tower crane performs the lifting task, expressed as: 5) Algorithm BC+PPO+GAIL process design; The behavior cloning algorithm is used to train the initial policy network by minimizing the difference between the actions generated by the strategy and the expert actions; in the model fine-tuning stage, the network parameter π of the behavior cloning strategy is used θ BC As the initial parameters of PPO and GAIL policy networks.

2. According to claim 1, a tower crane path intelligent planning method based on reinforcement learning and imitation learning is characterized in that: The specific steps of this method are as follows: Step 1: Through the investigation of the actual working environment of the tower crane driver, understand the space size of the tower crane cab, the layout of the joystick, the function of the buttons, and plan the operation area of ​​the tower crane driver; Step 2: Install the left and right positioning base stations of HTC VivePro to ensure accurate tracking of the head-mounted display; Step 3: Select Window->Assetstore in the top toolbar of Unity to open the UnityAsset store, download the SteamVR plug-in from it and install it into the Unity project; Step 4: Add the "CameraRig" component of SteamVR at the height of the crane driver's perspective in the virtual environment crane cab, set relevant parameters, and define the crane driver's perspective in the virtual environment, including the position and direction of the head; Step 5: Select Edit->ProjectSettings->InputManager in the top toolbar of Unity, add three new axes to the joystick, and name them "HookVertical", "HookHorizontal" and "Rotation" to correspond to the movement of the crane in three degrees of freedom; then configure the parameters of these axes to match the corresponding inputs of the joystick; set "Type" to "JoystickAxis", according to the function of the actual tower crane joystick, the left-hand joystick x-axis controls the rotation of the tower crane, the y-axis controls the amplitude change of the trolley, and the right-hand joystick y-axis controls the lifting of the hook, and the joystick axis is mapped to the corresponding InputManager axis; Step 6: Connect the power supply of the base station and link box accessories to ensure that the head mounted display is in normal working condition; Step 7: Connect the joystick and change the key control of the tower crane dynamic model to axis control; Step 8: Click Unity 3D to run the program; the operator performs an immersive tower crane simulation lifting task through a head-mounted display.

3. According to claim 1, a tower crane path intelligent planning method based on reinforcement learning and imitation learning is characterized in that: The crane's own status information is contained in the parameters of formula (1): composition: S crane =[θ t ,h t ,d t ,r t ,c t ,g t ,b t ,l t ] (1) Among them, θ t represents the rotation angle of the tower arm at time step t; h t represents the height of the tower crane hook from the ground at time step t; d t represents the radial distance between the tower crane's rotation center and the luffing trolley at time step t; r t Indicates the ratio of the current moving distance to the distance between the starting point and the target point at time step t; c t represents the direction vector from the hook position to the target point position at time step t; g t Represents the dot product value between the direction vector and the ground normal vector; b t is a binary value used to determine whether the lifting task is completed at time step t; l t Indicates the position information of the material in three-dimensional space at a certain moment.

4. According to claim 1, a tower crane path intelligent planning method based on reinforcement learning and imitation learning is characterized in that: The action space is represented as: A=[T t ,J t ,H t ] (2) T t , H t and J t is the motion vector of the tower crane, reflecting the kinematic characteristics of the tower crane, namely the luffing motion of the trolley, the lifting motion of the hook and the slewing motion of the tower arm; the motion of the tower crane is normalized and restricted to [-1,1] after the motion is selected; formula (3) represents the specific normalization method: Among them, A norm ∈[-1,1],A∈[A min ,A max ], A min and A max They represent the minimum and maximum values ​​of the corresponding elements in the action variable respectively.

5. The tower crane path intelligent planning method based on reinforcement learning and imitation learning according to claim 1 is characterized in that: The reward function in the tower crane agent reward function design is expressed in the form of formula (4): R1=clip(k distance (P t-1 -P t ),-0.1,+0.1) (4) Among them, P t-1 represents the distance between the tower crane hook and the target position at step t-1, P t represents the distance between the crane hook and the target position at time step t, k distance is a constant coefficient; the distance reward value of each time step is limited between -0.1 and 0.1 through the clip function; as shown in formula (5), the distance calculation formula between the hook and the target position is: Among them, (x hook ,y hook ,z hook ) is the current position of the hook in three-dimensional space, (x goal ,y goal ,z goal ) is the target point of the tower crane lifting task.

6. The tower crane path intelligent planning method based on reinforcement learning and imitation learning according to claim 1 is characterized in that: The tower arm heading reward function is expressed as: R2=clip(k direction (D t-1 -D t ),-0.1,+0.1) (6) Among them, D t-1 represents the angle between the tower arm and the target point in the xy plane at the t-1 time step, D t represents the angle between the tower arm and the target point in the xy plane at time step t, k direction is a constant coefficient; the distance reward value of each time step is limited to between -0.1 and 0.1 through the clip function.

7. The tower crane path intelligent planning method based on reinforcement learning and imitation learning according to claim 1 is characterized in that: The obstacles at the construction site are of two types: buildings and materials piled up on the construction site. The reward function also includes the collision judgment between the tower crane body and these two obstacles. In the actual tower crane lifting operation, based on safety considerations, a certain safety distance is required between the tower crane and obstacles. Taking into account the complexity of the virtual environment and the operating speed of the tower crane, the radius of the safety buffer zone is set to 5m. That is, when the distance between the tower crane and the obstacle is less than 5m, it is considered that a collision has occurred. At this time, the tower crane agent receives a discrete reward R3 of -2.

8. The tower crane path intelligent planning method based on reinforcement learning and imitation learning according to claim 1 is characterized in that: Update network parameters using gradient descent To minimize the loss function θ of behavior cloning; the parameter update formula is expressed as (8): In the model fine-tuning phase, the network parameters using the behavior cloning strategy As the initial parameters of the PPO and GAIL policy networks; define the weight coefficient λ(N) related to the number of iterations N to weigh the optimization of the two algorithms on the policy network. Formula (9) represents the expression of the hybrid loss function: L mix =λ(N)L GAIL (θ)+(1-λ(N))L PPO (f) (9) Among them, λ(N) is the weight coefficient, which is used to balance the losses of PPO and GAIL; L GAIL (θ) represents the loss function of GAIL, which focuses on imitating the effect of expert strategy; L PPO (φ) represents the loss function of PPO; During the training process, λ(N) decreases continuously as the number of iterations increases, and a linear decay strategy is used to reduce the weight coefficient.

9. The tower crane path intelligent planning method based on reinforcement learning and imitation learning according to claim 8 is characterized in that: In the linear attenuation strategy to reduce the weight coefficient, first set an initial weight λ start And the final desired weight λ goal ; Then, the weight decay of each iteration is calculated according to the total number of iterations N of model training. The weight decay of each iteration is shown in formula (10): In each iteration, the weight coefficient is updated according to the attenuation weight Δλ calculated by formula (11): λ'=max(λ-Δλ,λ goal ) (11) where λ' represents the new weight coefficient and λ represents the old weight coefficient; the max function is used to ensure that λ does not fall below the target value λ goal .

10. The tower crane path intelligent planning method based on reinforcement learning and imitation learning according to claim 9 is characterized in that: λ start Set to 0.8, λ goal It is set to 0.2 to ensure that the training transitions from imitation learning to trial-and-error learning by interacting with the environment.

Citation Information

Cited By

  • Intelligent crane anti-swing control method based on reinforcement learning

    CN120270909A

  • Intelligent tower crane hoisting control method based on reinforcement learning

    CN120534875A

  • Robot, robot control method and system and readable storage medium

    CN122165425A