Tower crane swing control method and system based on near-end strategy optimization algorithm and computer storage medium
Through the proximal strategy optimization algorithm combined with the tower crane dynamic model and state space model, the control torque is optimized in real time, solving the swing problem of the tower crane in complex environments, achieving stable load transportation and safety improvement.
Patent Information
- Application Number
- CN202510582629.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-26
AI Technical Summary
The existing tower crane swing control methods are difficult to adapt to complex and changeable environments, resulting in unstable load transportation and safety risks.
The proximal strategy optimization algorithm (PPO) is used to combine the tower crane dynamic model and state space model, and the sensor measures environmental factors in real time, and optimizes the control torque using the policy network and the value network to achieve real-time control of the tower crane.
Effectively suppress the swing of the tower crane, improve the stability and safety of load transportation, and improve operational efficiency.
Smart Images

Figure CN120534864A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to tower crane control, in particular to a tower crane swing control method, system and computer storage medium based on a proximal strategy optimization algorithm. Background Art
[0002] like Figure 1 As shown, when a tower crane is lifting an object, it is prone to lateral and longitudinal swings due to the movement of the load and environmental factors (such as wind speed, changes in load size, etc.), which affects the safety and efficiency of the operation. Traditional tower crane swing control methods include manual operating levers, mechanical limiters, and relay control systems, among which: 1) For manual operating levers, control is performed through human control and mechanical transmission, relying on manual experience, and the adjustment efficiency is low; 2) For mechanical limiters, physical blocks are forced to limit, and there is no intelligent early warning, resulting in poor adaptability to dynamic environments; 3) For relay control systems, the motor is controlled by an electromagnetic switch, with a slow response speed (>0.5 seconds) and inability to make high-frequency and precise adjustments. In short, the existing tower crane swing control methods are difficult to adapt to complex and changing environments, and the control effect on the tower crane is very limited, resulting in unstable load transportation and a large safety risk.
[0003] Given that reinforcement learning algorithms have shown significant advantages in the control field, especially the proximal policy optimization (PPO) algorithm, which has become an effective tool for solving complex control problems due to its stability and efficiency, it is expected to combine the PPO algorithm with environmental factors to design a new tower crane swing control method, which has important practical significance and application value in the engineering field. Summary of the Invention
[0004] The purpose of the present invention is to provide a tower crane swing control method, system and computer storage medium based on a proximal strategy optimization algorithm, which can effectively control the swing of the tower crane in a complex and changeable environment, thereby smoothly transporting the load and improving the safety of the tower crane operation.
[0005] The present invention is achieved through the following technical solutions:
[0006] A tower crane swing control method based on a proximal strategy optimization algorithm comprises the following steps:
[0007] Step 1: Establish a tower crane dynamics model including environmental factors, expressed as follows:
[0008]
[0009] Where: and θ y They are the horizontal swing angle and the longitudinal swing angle of the tower crane respectively; and is the lateral control torque and longitudinal control torque; m is the mass of the tower crane load; Δm is the change in the mass of the tower crane load; g is the acceleration of gravity, F w,x and F w,y are the lateral and longitudinal components of the wind speed; L is the length of the rope; and are the lateral swing angular velocity and longitudinal swing angular velocity of the tower crane respectively;
[0010] Step 2: Establish a state space model Expressed as an equation:
[0011]
[0012] Where: s is the state vector, u is the control input vector, A, B, C and d are all environmental disturbance vectors, and s, u, A, B, C and d are respectively expressed by matrices as:
[0013]
[0014]
[0015] Step 3: Initialize the policy network π θ and value network ∨ φ , and set the hyperparameters of the proximal policy optimization algorithm;
[0016] Step 4: Substitute s, Δm, and F w,x and F w,y Input the proximal strategy optimization algorithm and output the lateral control torque and longitudinal control torque Update the state vector of the tower crane according to the state space model Collecting crane trajectory data in: and The time steps are and The state vector of yes The control input vector at time t, including the lateral control torque and longitudinal control torque is the reward function value, expressed as:
[0017]
[0018] Where: α, β, γ and δ are weight coefficients; P t is the current position of the load, P target is the target position of the load, is the control input vector used to penalize control actions whose swing amplitude exceeds a threshold;
[0019] Step 5: Use the generalized advantage estimate GAE to calculate the advantage value at each time step Expressed as:
[0020]
[0021] Where: γ is the discount factor, λ is the hyperparameter of the generalized advantage estimator GAE, and Represents the time step and The state value of
[0022] Step 6: Calculate the clipping objective function L separately CLIP (θ), value function error L VF (φ) and the entropy regularization term Then update the policy network π θ The policy parameters θ and the value network ∨ φ The value function parameter φ;
[0023] Step 7: Repeat steps 4 to 6 until the proximal policy optimization algorithm converges to obtain a trained proximal policy optimization algorithm.
[0024] Step 8: Real-time collection θ y 、 Δm、F w,x and F w,y Input the trained proximal policy optimization algorithm and output the control input vector and state value Thereby controlling the swing of the tower crane in real time.
[0025] Furthermore, the hyperparameters of the proximal strategy optimization algorithm in step 3 include a clipping range ∈, a discount factor γ, and a generalized advantage estimation parameter λ.
[0026] Furthermore, in step 6, the objective function L is clipped CLIP (θ), value function error L VF (φ) and the entropy regularization term Respectively expressed as:
[0027]
[0028] Where: is the importance sampling ratio, ∈ is the cropping range, and θ is the policy network π θ The strategy parameters, The advantage value for each time step
[0029]
[0030] Where φ is the value network V φ The value function parameters, is the time step The predicted state value of is the time step The target state value of
[0031] Entropy regularization term Expressed as:
[0032]
[0033] Where, is a parameterized policy function, for Down-strategy network π θ The conditional probability of choosing action a, is the time step The state vector of .
[0034] Furthermore, the value of ∈ in step 6 is 0.1 to 0.2.
[0035] Furthermore, in step 6, the updated strategy network π θ The policy parameters θ and the value network V φ The process of value function parameter φ is: θ The policy parameter θ is updated as The value network V φ The value function parameter θ is updated as
[0036] A tower crane swing control system includes a sensor module, a control module and an execution module, wherein:
[0037] The sensor module is used to measure the lateral swing angle of the tower crane in real time Lateral angular velocity Longitudinal swing angle θ y , longitudinal oscillation angular velocity The lateral component of wind speed F w,x , the horizontal and vertical components of wind speed F w,y and the change in the crane load mass Δm;
[0038] The control module is used to execute a tower crane swing control method based on a proximal strategy optimization algorithm;
[0039] The execution module controls the input vector Adjusting the lateral control torque of the tower crane and longitudinal control torque
[0040] A computer storage medium stores a computer program. When the computer program is executed by a processor, a tower crane swing control method based on a proximal strategy optimization algorithm is implemented.
[0041] The present invention comprehensively considers the influence of various environmental factors on the swing angle of the tower crane, including wind speed, load mass change, and lateral and longitudinal swing angular velocity of the tower crane. By establishing a dynamic model of the tower crane swing, a state space model is derived, and then the state vector, load mass change and wind speed involved in the state space model are used as inputs of the proximal policy optimization algorithm (PPO), thereby collecting the trajectory data of the tower crane during the interaction between the current proximal policy optimization algorithm and the environment. Because the value reward function in the trajectory data combines the swing angle, the position error of the load and the control input vector, it can encourage the load to reach the target position quickly, stably and low. Therefore, the generalized advantage estimation (GAE) and the target loss function are combined to optimize the policy network and the value network, so that the PPO algorithm can adapt to environmental changes, thereby effectively suppressing the swing of the tower crane and realizing the smooth transportation of the load, which has very important practical significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is the force analysis diagram of the tower crane in working condition;
[0043] Figure 2 This is a flow chart of the tower crane swing control method of the present invention. DETAILED DESCRIPTION
[0044] The present invention will be further described in detail below with reference to specific embodiments, which are intended to explain the present invention rather than to limit it.
[0045] Example 1
[0046] refer to Figure 2 As shown, a tower crane swing control method based on a proximal strategy optimization algorithm includes the following steps:
[0047] Step 1: Establish a tower crane dynamics model including environmental factors, expressed as follows:
[0048]
[0049] Where: and θ y They are the horizontal swing angle and the longitudinal swing angle of the tower crane respectively; and is the lateral control torque and longitudinal control torque; m is the mass of the tower crane load; Δm is the change in the mass of the tower crane load; g is the acceleration of gravity, F w,x and F w,yare the lateral and longitudinal components of the wind speed; L is the length of the rope; and are the lateral and longitudinal swing angular velocities of the tower crane respectively;
[0050] Step 2: Establish a state space model Expressed as an equation:
[0051]
[0052] Where: s is the state vector, u is the control input vector, A, B, C and d are all environmental disturbance vectors, and s, u, A, B, C and d are respectively expressed by matrices as:
[0053]
[0054] Step 3: Initialize the policy network π θ and value network V φ , and set the hyperparameters of the proximal policy optimization algorithm, including the clipping range ∈, the discount factor γ, and the generalized advantage estimation parameter λ;
[0055] Step 4: Substitute s, Δm, and F w,x and F w,y Input the proximal strategy optimization algorithm and output the lateral control torque and longitudinal control torque Update the state vector of the tower crane according to the state space model Collecting crane trajectory data in: and The time steps are and The state vector of yes The control input vector at time t, including the lateral control torque and longitudinal control torque is the reward function value, which takes into account the swing angle, load position error and control input vector to encourage suppressing the swing and reaching the target position quickly, expressed as:
[0056]
[0057] Where: α, β, γ and δ are weight coefficients; and θ y are the horizontal swing angle and the longitudinal swing angle of the tower crane respectively; P t is the current position of the load, P target is the target position of the load, is the control input vector used to penalize control actions whose swing amplitude exceeds a threshold;
[0058] Step 5: Use the generalized advantage estimate GAE to calculate the advantage value at each time step Expressed as:
[0059]
[0060] Where: γ is the discount factor, λ is the hyperparameter of the generalized advantage estimator GAE, and Represents the time step and The state value of
[0061] Step 6: Calculate the clipping objective function L separately CLIP (θ), value function error L VF (φ) and the entropy regularization term Respectively expressed as:
[0062]
[0063] Where: is the importance sampling ratio; ∈ is the clipping range, which ranges from 0.1 to 0.2; θ is the policy network π θ Strategy parameters; The advantage value for each time step Limiting the policy network π through pruning θ The update amplitude makes the training process more stable;
[0064]
[0065] Where φ is the value network V φ The value function parameters, is the time step The predicted state value of is the time step The target state value of
[0066] Entropy regularization term Expressed as:
[0067]
[0068] Where, is a parameterized policy function, for Down-strategy network π θ The conditional probability of choosing action a, is the time step The state vector of
[0069] Then, the policy network πθ The policy parameter θ is updated as The value network V φ The value function parameter θ is updated as
[0070] Where η is the learning rate, The change in the policy parameter θ, is the change in the value function parameter θ;
[0071] Step 7: Repeat steps 4 to 6 until the proximal policy optimization algorithm converges to obtain a trained proximal policy optimization algorithm.
[0072] Step 8: Collect the real-time θ y 、 Δm、F w,x and F w,y Input the trained proximal policy optimization algorithm and output the control input vector and state value Thereby controlling the swing of the tower crane in real time.
[0073] This embodiment addresses the problem of lateral and longitudinal swing caused by external factors (such as wind speed, load weight changes, and load movement) during the execution of a tower crane task. Based on the proximal policy optimization (PPO) algorithm, this embodiment introduces environmental factors and proposes a tower crane swing control method based on the proximal policy optimization (PPO) algorithm. The method suppresses the tower crane swing through efficient strategy optimization capabilities and dynamic adaptability to the environment, thereby improving operation safety and efficiency.
[0074] Example 2
[0075] Based on the crane swing control method proposed in Example 1, this embodiment proposes a tower crane swing control system, including a sensor module, a control module, and an execution module, wherein:
[0076] The sensor module is used to measure the lateral swing angle of the tower crane in real time Lateral angular velocity Longitudinal swing angle θ y , longitudinal oscillation angular velocity The lateral component of wind speed F w,x , the horizontal and vertical components of wind speed F w,y and the change in the crane load mass Δm;
[0077] The control module is used to execute Figure 1 The tower crane swing control method based on the proximal strategy optimization algorithm shown in FIG;
[0078] The execution module controls the input vector Adjusting the lateral control torque of the tower crane and longitudinal control torque
[0079] Example 3
[0080] Based on the pendulum swing control method proposed in Example 1, this embodiment proposes a computer storage medium storing a computer program. When the computer program is executed by a processor, the following is achieved: Figure 1 The tower crane swing control method based on the proximal strategy optimization algorithm is shown.
Claims
1. A tower crane swing control method based on proximal strategy optimization algorithm, characterized in that: The following steps are involved: Step 1: Establish a tower crane dynamics model including environmental factors, expressed as follows: Where: θ x and θ y They are the horizontal swing angle and the longitudinal swing angle of the tower crane respectively; and is the lateral control torque and longitudinal control torque; m is the mass of the tower crane load; Δm is the change in the mass of the tower crane load; g is the acceleration of gravity, F w,x and F w,y are the lateral and longitudinal components of the wind speed; L is the length of the rope; and are the lateral swing angular velocity and longitudinal swing angular velocity of the tower crane respectively; Step 2: Establish a state space model Expressed as an equation: Where: s is the state vector, u is the control input vector, A, B, C and d are all environmental disturbance vectors, and s, u, A, B, C and d are respectively expressed by matrices as: Step 3: Initialize the policy network π θ and value network V φ , and set the hyperparameters of the proximal policy optimization algorithm; Step 4: Substitute s, Δm, and F w,x and F w,y Input the proximal strategy optimization algorithm and output the lateral control torque and longitudinal control torque Update the state vector of the tower crane according to the state space model Collecting crane trajectory data in: and The time steps are and The state vector of yes The control input vector at time t, including the lateral control torque and longitudinal control torque is the reward function value, expressed as: Where: α, β, γ and δ are weight coefficients; P t is the current position of the load, P target is the target position of the load, is the control input vector used to penalize control actions whose swing amplitude exceeds a threshold; Step 5: Use the generalized advantage estimate GAE to calculate the advantage value at each time step Expressed as: Where: γ is the discount factor, λ is the hyperparameter of the generalized advantage estimator GAE, and Represents the time step and The state value of Step 6: Calculate the clipping objective function L separately CLIP (θ), value function error L VF (φ) and the entropy regularization term Then update the policy network π θ The policy parameters θ and the value network V φ The value function parameter φ; Step 7: Repeat steps 4 to 6 until the proximal policy optimization algorithm converges to obtain a trained proximal policy optimization algorithm. Step 8: Real-time collection θ y 、 Δm、F w,x and F w,y Input the trained proximal policy optimization algorithm and output the control input vector and state value Thereby controlling the swing of the tower crane in real time.
2. The tower crane swing control method based on the proximal strategy optimization algorithm according to claim 1 is characterized in that: The hyperparameters of the proximal strategy optimization algorithm in step 3 include the clipping range ∈, the discount factor γ, and the generalized advantage estimation parameter λ.
3. The tower crane swing control method based on the proximal strategy optimization algorithm according to claim 1 is characterized in that: In step 6, the objective function L is cut CLIP (θ), value function error L VF (φ) and the entropy regularization term Respectively expressed as: Where: is the importance sampling ratio, ∈ is the cropping range, and θ is the policy network π θ The strategy parameters, The advantage value A for each time step t ; Where φ is the value network V φ The value function parameters, is the time step The predicted state value of is the time step The target state value of Entropy regularization term Expressed as: Where, is a parameterized policy function, for Down-strategy network π θ The conditional probability of choosing action a, is the time step The state vector of .
4. The tower crane swing control method based on the proximal strategy optimization algorithm according to claim 3 is characterized in that: The value of ∈ in step 6 is 0.1 to 0.
2.
5. The tower crane swing control method based on the proximal strategy optimization algorithm according to claim 3 is characterized in that: In step 6, the policy network π is updated θ The policy parameters θ and the value network V φ The process of value function parameter φ is: θ The policy parameter θ is updated as The value network V φ The value function parameter θ is updated as 6. A tower crane swing control system, characterized in that: It includes sensor module, control module and execution module, among which: The sensor module is used to measure the lateral swing angle of the tower crane in real time Lateral angular velocity Longitudinal swing angle θ y , longitudinal oscillation angular velocity The lateral component of wind speed F w,x , the horizontal and vertical components of wind speed F w,y and the change in the crane load mass Δm; The control module is used to execute the tower crane swing control method based on the proximal strategy optimization algorithm according to any one of claims 1 to 5; The execution module controls the input vector Adjusting the lateral control torque of the tower crane and longitudinal control torque 7. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a processor, the tower crane swing control method based on the proximal strategy optimization algorithm as described in any one of claims 1 to 5 is implemented.
Citation Information
Cited By
Automatic deviation rectification control method for tower crane moment limiter
CN120736414A
PINN-based tower crane motion trail quality evaluation method, apparatus and device, and medium
CN121327370A