Suspension bridge prefabricated cable strand method construction traction system control method based on reinforcement learning
By optimizing PID parameters through a multi-objective reward function based on reinforcement learning and the DDPG algorithm, the problems of low control accuracy and efficiency in the construction of suspension bridges using the prefabricated cable strand method were solved, and the intelligent construction and stability of the prefabricated cable strand method of suspension bridges were improved.
Patent Information
- Application Number
- CN202510652835.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing prefabricated cable strand construction method for suspension bridges, the PID control algorithm has a lag in manual parameter adjustment, resulting in low traction speed control accuracy, low efficiency and insufficient intelligence, making it difficult to adapt to the complex and changing construction environment and high-quality construction requirements.
A reinforcement learning-based approach is used to construct a multi-objective reward function and the DDPG algorithm to dynamically optimize PID control parameters. Through Simulink modeling and sensor data feedback, precise control of the tractor's operating speed is achieved. The traction system control is optimized by combining the state space of the tractor's position, velocity, and acceleration with the action space of the PID controller parameters.
It significantly improves the control accuracy and efficiency of the prefabricated cable strand construction method for suspension bridges, reduces the need for human intervention, and achieves stable, efficient, and intelligent control of the cable strand traction process.
Smart Images

Figure CN120704108A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bridge construction, and in particular relates to a reinforcement learning-based construction traction system control method for a suspension bridge using a prefabricated cable strand method. Background Art
[0002] Currently, when using the PPWS method for erecting the main cable strands of a suspension bridge, the classic PID algorithm is the primary control method for the traction winch. The PID (Proportional-Integral-Derivative) control algorithm is a widely used automatic control algorithm in industrial control. It adjusts the control variable based on the system error (i.e., the difference between the set value and the actual value), thereby achieving system control.
[0003] However, as bridge spans gradually increase, existing control theories are unable to adapt to the complex and changing construction environment and high-quality construction requirements. Due to the lag in manual parameter adjustment, there are problems such as low traction speed control accuracy, low traction efficiency, and insufficient intelligence level. It is urgent to propose an intelligent control method suitable for PPWS cable traction to ensure the stability, efficiency, and intelligent controllability of the cable traction process. Summary of the Invention
[0004] In response to the problems of large puller speed fluctuations, insufficient trajectory smoothness and poor stability of the traction process caused by the lag in manual parameter adjustment in the existing construction control of prefabricated cable-strand method of suspension bridges, the purpose of the present invention is to provide a control method for the traction system of the prefabricated cable-strand method of suspension bridge construction based on reinforcement learning. By constructing a multi-objective reinforcement learning framework that includes trajectory smoothness rewards, speed stability rewards and speed mutation penalties, combined with the DDPG algorithm to dynamically optimize PID control parameters, precise regulation of the puller operating speed is achieved, significantly improving the control accuracy and construction efficiency of the main cable traction process, while reducing the need for human intervention.
[0005] The technical solution of the present invention is: a method for controlling a traction system for construction of a suspension bridge prefabricated cable strand method based on reinforcement learning, comprising the following steps: S1: Modeling the PPWS cable pulling process using Simulink; S2: Build a reinforcement learning environment and define the state space and action space, where the state space includes the position, velocity, and acceleration of the dragger, and the action space includes the PID controller parameters, specifically the reward function including trajectory smoothness, velocity stability, and velocity mutation penalty; S3: Use the DDPG reinforcement learning model for training, initialize the neural network, and set training parameters; S4: Deploy the trained DDPG strategy network to the hardware control system, dynamically optimize the PID parameters based on the status data collected by the sensors, and achieve intelligent control of the cable strand traction.
[0006] In step S1, Simulink is used to model the traction winch, puller, pulley block, counterweight, workload and traction cable of the traction system of the prefabricated cable strand method for the suspension bridge, and a simulation environment for the main cable traction process is established.
[0007] In step S2, a reinforcement learning environment is built, and the construction process of the PPWS method is combined to define the state space s (x, z, v, a), where x and z are the horizontal coordinates and elevation coordinates of the puller, v is the actual speed, and a is the acceleration of the puller; the action space at (Kp, Ki, Kd), R total Reward function and sampling interval, Kp, Ki, Kd are the proportional coefficient, integral coefficient and differential coefficient of the traction equipment respectively.
[0008] The reward function R in step S2 total It is the feedback of the environment to the agent's actions, which is used to quantify the quality of the action selection. The reward function R total The setting takes into account the running speed and stability of the puller, and includes the following parts: R total =R smooth +R acc +R vel (1) (2) (3) (4) Among them, R total is the total reward score, R smooth is the trajectory smoothness reward, R vel is the speed stability reward, R acc is the speed mutation penalty, k1, k2, c1, c2 are their respective coefficients, calibrated according to the model training effect, v min is the minimum allowable speed, v max To allow the maximum speed, set according to the cable traction construction requirements and winch performance; v design is the expected running speed; dt is the sampling interval, N is the total number of time steps, which represents the time window length of the dragger trajectory monitoring; t is the time step number; kt is the trajectory curvature of the dragger at time t, and the absolute difference |k t -k t-1 |Measures the change in trajectory curvature between adjacent time steps. A larger difference indicates a less smooth trajectory.
[0009] In step S3, the DDPG reinforcement learning model is used for training, the neural network is initialized, and the training parameters are set. The specific steps are as follows: S31: Initialize the network structure and create four neural networks: the online actor network (policy network, used to generate deterministic actions), the target actor network (used for stable training, parameters synchronized with the online actor network), the online critic network (Q value network, used to evaluate the value of state-action pairs), and the target critic network (calculate target Q values, parameters synchronized with the online critic network). The experience replay pool is also initialized to store interaction samples. S32: Set the number of training rounds and the number of steps for each round of training, set the number of batches for each group of model training (batch_size), and give the initial state S of the puller t (x t , z t , v t , a t ); S33: Agent training, the specific process is as follows: S331: Taking state S as the input of the model, the online actor network generates a deterministic action A t (Kp t 、Ki t 、Kd t ), and add exploration noise; S332: According to Kp after adding noise t 、Ki t 、Kd t The deviation Δv between the actual speed and the expected speed is used for PID calculation to calculate the control parameters sent to the traction winch motor; S333: The Simulink model updates the state space according to the motor operating parameters to obtain S t+1 , and calculate the single-step reward r t , the data (S t 、A t 、r t 、S t+1 ) is stored in the experience replay pool. If the number of samples in the experience replay pool is less than the set batch number, steps S331-S333 are repeated; S334: Train and update the network, specifically including: Critics network update: Randomly sample a batch of samples from the experience replay pool, the target actor generates the target action based on the next state, the target critic calculates the target Q value, calculates the critic's loss function (mean square error), and updates the parameters of the online critic network through gradient descent; Actor network update: The goal of the online actor is to maximize the Q value and update the parameters of the online actor network through gradient ascent; Adopt a soft update strategy to slowly synchronize target network parameters; S335: Repeat S331-S334. When the network converges or reaches the preset number of iterations, training is started. The online actor network (strategy model) is the final training result.
[0010] The technical effects of the present invention are: 1. The control method of the traction system of the prefabricated cable method for suspension bridge construction based on reinforcement learning can automatically adjust the values of PID parameters according to the speed and stability of the puller operation, effectively solving the problem of poor parameter setting flexibility in traditional PID control, significantly improving the control accuracy and stability of cable traction, and improving construction efficiency; 2. The present invention defines a comprehensive state space including the position, speed and acceleration of the puller, and an action space including the proportional, integral and differential coefficients of the PID controller, providing clear input and output for the reinforcement learning model, so that the model can more accurately perceive the construction status and make control decisions; 3. The present invention designs a multi-objective reward function including trajectory smoothness, speed stability and speed mutation penalty, so that the model can simultaneously consider multiple performance indicators during the training process, balancing the requirements of construction accuracy, efficiency and stability.
[0011] The following is a further description with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a flow chart of intelligent control of cable strand traction according to the present invention. DETAILED DESCRIPTION Example 1
[0013] The reinforcement learning-based control method for the traction system of the prefabricated cable strand method for suspension bridge construction includes the following steps: S1: Modeling the PPWS cable pulling process using Simulink; S2: Build a reinforcement learning environment and define the state space and action space, where the state space includes the position, velocity, and acceleration of the dragger, and the action space includes the PID controller parameters, specifically the reward function including trajectory smoothness, velocity stability, and velocity mutation penalty; S3: Use the DDPG reinforcement learning model for training, initialize the neural network, and set training parameters; S4: Deploy the trained DDPG strategy network to the hardware control system, dynamically optimize the PID parameters based on the status data collected by the sensors, and achieve intelligent control of the cable strand traction.
[0014] In step S1, Simulink is used to model the traction winch, puller, pulley block, counterweight, workload and traction cable of the traction system of the prefabricated cable strand method for the suspension bridge, and a simulation environment for the main cable traction process is established.
[0015] In step S2, a reinforcement learning environment is built, and the construction process of the PPWS method is combined to define the state space s (x, z, v, a), where x and z are the horizontal coordinates and elevation coordinates of the puller, v is the actual speed, and a is the acceleration of the puller; the action space at (Kp, Ki, Kd), R total Reward function and sampling interval, Kp, Ki, Kd are the proportional coefficient, integral coefficient and differential coefficient of the traction equipment respectively.
[0016] The reward function R in step S2 total It is the feedback of the environment to the agent's actions, which is used to quantify the quality of the action selection. The reward function R total The setting takes into account the running speed and stability of the puller, and includes the following parts: R total =R smooth +R acc +R vel (1) (2) (3) (4) Among them, R total is the total reward score, R smooth is the trajectory smoothness reward, R vel is the speed stability reward, R acc is the speed mutation penalty, k1, k2, c1, c2 are their respective coefficients, calibrated according to the model training effect, v min is the minimum allowable speed, v max To allow the maximum speed, set according to the cable traction construction requirements and winch performance; v design is the expected running speed; dt is the sampling interval, N is the total number of time steps, which represents the time window length of the dragger trajectory monitoring; t is the time step number; kt is the trajectory curvature of the dragger at time t, and the absolute difference |k t -k t-1 |Measures the change in trajectory curvature between adjacent time steps. A larger difference indicates a less smooth trajectory.
[0017] In step S3, the DDPG reinforcement learning model is used for training, the neural network is initialized, and the training parameters are set. The specific steps are as follows: S31: Initialize the network structure and create four neural networks: the online actor network (policy network, used to generate deterministic actions), the target actor network (used for stable training, parameters synchronized with the online actor network), the online critic network (Q value network, used to evaluate the value of state-action pairs), and the target critic network (calculate target Q values, parameters synchronized with the online critic network). The experience replay pool is also initialized to store interaction samples. S32: Set the number of training rounds and the number of steps for each round of training, set the number of batches for each group of model training (batch_size), and give the initial state S of the puller t (x t , z t , v t , a t ); S33: Agent training, the specific process is as follows: S331: Taking state S as the input of the model, the online actor network generates a deterministic action A t (Kp t 、Ki t 、Kd t ), and add exploration noise; S332: According to Kp after adding noise t 、Ki t 、Kd t The deviation Δv between the actual speed and the expected speed is used for PID calculation to calculate the control parameters sent to the traction winch motor; S333: The Simulink model updates the state space according to the motor operating parameters to obtain S t+1 , and calculate the single-step reward r t , the data (S t 、A t 、r t 、S t+1 ) is stored in the experience replay pool. If the number of samples in the experience replay pool is less than the set batch number, steps S331-S333 are repeated; S334: Train and update the network, specifically including: Critics network update: Randomly sample a batch of samples from the experience replay pool, the target actor generates the target action based on the next state, the target critic calculates the target Q value, calculates the critic's loss function (mean square error), and updates the parameters of the online critic network through gradient descent; Actor network update: The goal of the online actor is to maximize the Q value and update the parameters of the online actor network through gradient ascent; Adopt a soft update strategy to slowly synchronize target network parameters; S335: Repeat S331-S334. When the network converges or reaches the preset number of iterations, training is started. The online actor network (strategy model) is the final training result.
[0018] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.
Claims
1. A reinforcement learning-based control method for the traction system of a suspension bridge prefabricated cable strand construction, characterized by: The following steps are involved: S1: Modeling the PPWS cable pulling process using Simulink; S2: Build a reinforcement learning environment and define the state space and action space, where the state space includes the position, velocity, and acceleration of the dragger, and the action space includes the PID controller parameters, specifically the reward function including trajectory smoothness, velocity stability, and velocity mutation penalty; S3: Use the DDPG reinforcement learning model for training, initialize the neural network, and set training parameters; S4: Deploy the trained DDPG strategy network to the hardware control system, dynamically optimize the PID parameters based on the status data collected by the sensors, and achieve intelligent control of the cable strand traction.
2. The reinforcement learning-based control method for the traction system of a suspension bridge prefabricated cable strand construction method according to claim 1 is characterized by: In step S1, Simulink is used to model the traction winch, puller, pulley block, counterweight, workload and traction cable of the traction system of the prefabricated cable strand method for the suspension bridge, and a simulation environment for the main cable traction process is established.
3. The reinforcement learning-based control method for the traction system of a suspension bridge prefabricated cable strand construction method according to claim 2 is characterized by: In step S2, a reinforcement learning environment is built, and the construction process of the PPWS method is combined to define the state space s (x, z, v, a), where x and z are the horizontal coordinates and elevation coordinates of the puller, v is the actual speed, and a is the acceleration of the puller; the action space at (Kp, Ki, Kd), R total Reward function and sampling interval, Kp, Ki, Kd are the proportional coefficient, integral coefficient and differential coefficient of the traction equipment respectively.
4. The reinforcement learning-based control method for the traction system of a suspension bridge prefabricated cable strand construction method according to claim 3 is characterized by: The reward function R in step S2 total It is the feedback of the environment to the agent's actions, which is used to quantify the quality of the action selection. The reward function R total The setting takes into account the running speed and stability of the puller, and includes the following parts: R total =R smooth +R acc +R vel (1) (2) (3) (4) Among them, R total is the total reward score, R smooth is the trajectory smoothness reward, R vel is the speed stability reward, R acc is the speed mutation penalty, k1, k2, c1, c2 are their respective coefficients, calibrated according to the model training effect, v min is the minimum allowable speed, v max To allow the maximum speed, set according to the cable traction construction requirements and winch performance; v design is the expected running speed; dt is the sampling interval, N is the total number of time steps, which represents the time window length of the dragger trajectory monitoring; t is the time step number; kt is the trajectory curvature of the dragger at time t, and the absolute difference |k t -k t-1 |Measures the change in trajectory curvature between adjacent time steps. A larger difference indicates a less smooth trajectory.
5. The reinforcement learning-based control method for the traction system of a suspension bridge prefabricated cable strand construction method according to claim 1 is characterized by: In step S3, the DDPG reinforcement learning model is used for training, the neural network is initialized, and the training parameters are set. The specific steps are as follows: S31: Initialize the network structure and create four neural networks: the online actor network (policy network, used to generate deterministic actions), the target actor network (used for stable training, parameters synchronized with the online actor network), the online critic network (Q value network, used to evaluate the value of state-action pairs), and the target critic network (calculate target Q values, parameters synchronized with the online critic network). The experience replay pool is also initialized to store interaction samples. S32: Set the number of training rounds and the number of steps for each round of training, set the number of batches for each group of model training (batch_size), and give the initial state S of the puller t (x t , z t , v t , a t ); S33: Agent training, the specific process is as follows: S331: Taking state S as the input of the model, the online actor network generates a deterministic action A t (Kp t 、Ki t 、Kd t ), and add exploration noise; S332: According to Kp after adding noise t 、Ki t 、Kd t The deviation Δv between the observed actual speed and the expected speed is used for PID calculation to calculate the control parameters sent to the traction winch motor; S333: The Simulink model updates the state space according to the motor operating parameters to obtain S t+1 , and calculate the single-step reward r t , the data (S t 、A t 、r t 、S t+1 ) is stored in the experience replay pool. If the number of samples in the experience replay pool is less than the set batch number, steps S331-S333 are repeated; S334: Train and update the network, specifically including: Critics network update: Randomly sample a batch of samples from the experience replay pool, the target actor generates the target action based on the next state, the target critic calculates the target Q value, calculates the critic's loss function (mean square error), and updates the parameters of the online critic network through gradient descent; Actor network update: The goal of the online actor is to maximize the Q value and update the parameters of the online actor network through gradient ascent; Adopt a soft update strategy to slowly synchronize target network parameters; S335: Repeat S331-S334. When the network converges or reaches the preset number of iterations, training is started. The online actor network (strategy model) is the final training result.
Citation Information
Patent Citations
Method for determining unstressed length of branch cable strands of main cable of suspension bridge
CN113089452A
Suspension bridge main cable strand traction control system and method
CN113774806A
Bridge detection unmanned aerial vehicle autonomous navigation and stability control method based on reinforcement learning
CN113821044A
PID controller parameter self-tuning method based on reinforcement learning algorithm
CN118244618A
Suspension bridge air spinning method construction traction system control method based on reinforcement learning
CN120722719A
Cited By
Method for obtaining PID parameter setting intelligent agent and related device
CN121879095A
A method for obtaining a PID parameter setting intelligent agent and a related device
CN121879095B