Double-pendulum bridge crane self-adaptive control method and system based on deep reinforcement learning

By optimizing the PID controller parameters through deep reinforcement learning, the problem of the single network structure of the adaptive algorithm for double pendulum bridge cranes is solved, enabling rapid positioning and anti-swaying, and improving the dynamic performance and robustness of the system.

CN121523050APending Publication Date: 2026-02-13LANZHOU UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511871322.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

The existing adaptive algorithm network structure of double pendulum bridge cranes is simple and lacks the ability to explore complex nonlinear dynamic characteristics, resulting in insufficient positioning and swing suppression performance. Furthermore, it does not fully consider the coupling effect between the hook and the cargo, and the dynamic performance and stability of the system need to be improved.

Method used

An adaptive control method based on deep reinforcement learning is adopted. By establishing a mathematical model of the double pendulum bridge crane, combining a PID controller and a reinforcement learning algorithm, the PID controller parameters are optimized. The Euler-Lagrange method is used to consider the double pendulum effect and damping, and a performance index function is constructed to achieve rapid positioning and anti-swaying of the trolley.

Benefits of technology

It enables rapid positioning and effective anti-swaying of the double-swing bridge crane, improves the system's adaptability to parameter perturbations and external disturbances, and enhances the system's dynamic performance and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121523050A_ABST
    Figure CN121523050A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bridge crane control, in particular to a deep reinforcement learning-based self-adaptive control method and system for a double-pendulum bridge crane, and the system comprises a model building unit which is used for carrying out the modeling of the friction of a trolley, the damping of a lifting hook and the damping of goods in the movement process of the double-pendulum bridge crane, establishing a mathematical model of the double-pendulum bridge crane; the target setting unit is used for setting target position data of a trolley in the double-pendulum bridge crane and respectively inputting the target position data into the PID controller and the reinforcement learning algorithm; a PID (Proportion Integration Differentiation) controller of the crane driving unit receives the target position data and outputs driving force of a trolley to a mathematical model and a reinforcement learning algorithm of the double-pendulum bridge crane to drive the double-pendulum bridge crane to move; a data optimization unit, an error calculation unit and an error judgment unit; according to the self-adaptive control method, rapid positioning and effective swing prevention can be achieved, and meanwhile the adaptability of the system to parameter perturbation and external disturbance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bridge crane control technology, and in particular to an adaptive control method and system for a double pendulum bridge crane based on deep reinforcement learning. Background Technology

[0002] Bridge cranes, as important cargo transportation equipment, are widely used in engineering scenarios such as industrial production, construction sites, warehouse scheduling, and port trade. During crane operation, precise control of the trolley is required to reach the designated position. However, external disturbances and the movement of the trolley inevitably cause the hook and cargo to sway, which significantly reduces transportation efficiency and can even lead to safety accidents. Therefore, achieving good positioning, anti-sway control, and anti-interference performance are key issues in crane transportation.

[0003] In recent years, several automatic control algorithms, such as PID control, sliding mode control, predictive control, and active disturbance rejection control, have been studied and applied to the positioning and anti-sway control of cranes. Among them, PID control is widely used in intelligent crane control systems due to its simple structure, convenient engineering deployment, and good adaptability.

[0004] Among them, the article (Transportation control of double-pendulum cranes with anonlinear quasi-PID scheme: Design and experiments, 2018) proposes a PID-like coupled bridge crane control method with input constraints, which can guarantee accurate positioning and sway suppression under the influence of constraints and uncertainties.

[0005] Furthermore, the article (An enhanced coupling PD with sliding mode control method for underactuated double-pendulum overhead crane systems, 2019) proposes an enhanced coupled PD control method that combines sliding mode control, which reduces the impact of uncertainties and other disturbances on system performance to a certain extent.

[0006] Furthermore, patent (CN202210592475.9) constructs a crane speed-loop PID controller to achieve trolley speed tracking of the desired trajectory while ensuring that cargo swaying remains within a safe range. It is well known that the performance of these closed-loop control systems largely depends on the tuning of the PID parameters. However, in the aforementioned study, these parameters were pre-trained and have fixed values, making them unable to adapt to dynamic scene changes and limiting the performance of the control system to some extent.

[0007] To further improve the adaptability of crane systems under changing scenarios, developing an adaptive optimization algorithm for PID parameters is essential. Patent (CN202110641182.0) proposes a fuzzy adaptive PID control method for crane systems. Compared to traditional PID control, this method results in faster system adjustment, smaller overshoot, and stronger robustness. The paper (Proportional–integral-derivative controller with inlet derivative filter fine-tuning of a double-pendulum gantry crane system by a multi-objective genetic algorithm, 2020) proposes a multi-objective genetic algorithm to optimize the parameters of multi-loop PID control, which improves transportation efficiency and safety to some extent. Furthermore, patents (CN201711492808.6, CN202410653078.7) utilize neural networks as online estimators to optimize PID controller parameters in real time, demonstrating stronger adaptability to external disturbances.

[0008] The research on these adaptive control algorithms has proven to improve the response speed and adaptability of crane systems. However, some problems still remain to be solved: 1) Adaptive mechanisms based on fuzzy theory and neural networks rely too heavily on pre-set rules and network parameters, and still lack sufficient flexibility when facing unknown scenarios; 2) Most current studies do not consider the double-pendulum effect of cranes. In actual engineering, the mass of the hook cannot be ignored, and there is a coupling effect between it and the cargo, which reduces the control performance of the system; 3) The positioning and sway suppression performance needs improvement. The adaptive algorithm network structure in current research is relatively simple and lacks full exploration of complex nonlinear dynamic characteristics. The dynamic performance and stability of the system still need to be further improved. Summary of the Invention

[0009] This invention provides an adaptive control method for a double pendulum bridge crane based on deep reinforcement learning, which overcomes the shortcomings of the prior art. It can effectively solve the problem that the adaptive algorithm network structure used in existing double pendulum bridge cranes is too simple and cannot achieve rapid positioning of the trolley.

[0010] To address the above problems, one of the technical solutions of this invention is implemented through the following method: an adaptive control method for a double-swing bridge crane based on deep reinforcement learning, comprising the following steps: A mathematical model of the double pendulum bridge crane is established by modeling the friction of the trolley, the damping of the hook and the cargo during the movement of the double pendulum bridge crane. Set the target position data of the trolley in the double pendulum bridge crane and input it into the PID controller and reinforcement learning algorithm respectively; The PID controller receives the target position data and outputs the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the movement of the double pendulum bridge crane. Data on the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo during the movement of the double pendulum bridge crane are collected and output to the reinforcement learning algorithm. The reinforcement learning algorithm is trained using this data, combined with the target position data of the trolley and the driving force of the trolley, to optimize and update the parameters of the PID controller. The optimized parameters are then output to the PID controller. The displacement data of the trolley during the movement of the double pendulum bridge crane is output to the PID controller. After receiving the optimized parameters, the PID controller compares the received displacement data of the trolley with the target position data of the trolley and calculates the error between the displacement data and the target position data of the trolley. The system determines whether the error is zero. If no, the PID controller continues to output the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the double pendulum bridge crane to move. The above process is repeated until the error between the displacement data of the trolley and the target position data of the trolley is zero, which means that the trolley has reached the target position.

[0011] The above-mentioned modeling of the friction of the trolley, the damping of the hook and the cargo during the movement of the double pendulum bridge crane is carried out to establish a mathematical model of the double pendulum bridge crane, including: By applying the Euler-Lagrange method, the dynamic equations of the double-pendulum bridge crane are obtained as follows: , In the formula, For the mass of the car, For the mass of the hook, For the quality of the goods, This refers to the length of the rope connecting the trolley and the hook. The length of the rope connecting the hook and the cargo. For the displacement of the car, The swing angle of the hook, The swing angle of the goods. , and For the system's damping parameters, Represents gravitational acceleration. This indicates the driving force of the car.

[0012] The aforementioned PID controller receives target position data and outputs the trolley's driving force to the double-swing bridge crane's mathematical model and reinforcement learning algorithm, including: The driving force of the car at the current moment is calculated as follows: , In the formula, Let be the driving force of the car at time t; For small car The driving force of every moment; The change in driving force between two time points. The calculation method is as follows: , In the formula, This represents the parameters of the PID controller at time t. This represents the deviation between the target position and the actual displacement of the t vehicle at time t. ,in, This indicates the target position of the vehicle. This represents the displacement of the t vehicle at time t.

[0013] The aforementioned data on the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo during the movement of the double-swing bridge crane are collected and output to the reinforcement learning algorithm. This data, combined with the trolley's target position data and driving force, is used to train the reinforcement learning algorithm, optimizing and updating the parameters of the PID controller, including: Step 1: Set the current status of the double swing bridge crane system. The input is fed into the Actor network of the reinforcement learning algorithm, and the Actor network determines the input based on the current state. Calculate the actual action ; Step 2: The Actor network performs the actual actions. It interacts with the double-swing bridge crane to obtain a reward function. New Status of Double Swing Bridge Crane Systems ; Step 3: Put The data is placed into the experience replay pool. When the experience data in the experience replay pool reaches a certain amount, the initial data is deleted, and the experience data is fixed and saved to the replay pool. Step 4: Use importance sampling weighting techniques to extract a batch of empirical data from the experience replay pool, i.e., the historical batch. Calculate the values ​​of the two Critic networks and compare them, select the smaller value and update the Critic network; Step 5: Update the target Critic network and Actor network, then return to Step 1 and repeat this process until the set number of training iterations or performance metric values ​​are reached.

[0014] The above describes the current status of the double-swing bridge crane system. The input is fed into the Actor network of the reinforcement learning algorithm, and the Actor network determines the input based on the current state. Calculate the actual action ,include: Based on the agent's current state Computing the optimal strategy of reinforcement learning algorithms and set The optimal action output value is calculated as follows: , In the formula, As a strategy, This represents the optimal value of the strategy. This indicates that the function The largest value; Represents the total state variables of the double-swing bridge crane system. ; This represents the actual action output by the reinforcement learning algorithm. , Let t be the strategy of the action network in the reinforcement learning algorithm; To reinforce the learning algorithm policy in the state Entropy below; These are the weighting coefficients.

[0015] The aforementioned Actor network performs actual actions. It interacts with the double-swing bridge crane to obtain a reward function. New Status of Double Swing Bridge Crane Systems ,include: Define a performance metric function, i.e., a reward function. The calculation method is as follows: , In the formula, This indicates the deviation between the target position and the actual displacement of the trolley; The change in driving force between two time points; , and They are respectively , and The change in quantity.

[0016] The above-mentioned importance sampling weighting technique is used to extract a batch of empirical data from the experience replay pool, i.e., the historical batch. Calculate the values ​​of the two Critic networks and compare them, selecting the smaller value and updating the Critic network; including: Calculate the optimal state-value function in reinforcement learning algorithms and action value function The calculation method is as follows: , , In the formula, This is the discount factor. Indicates the strategy The following expectations; Calculate the value of the Critic network, i.e., the action-value function. Further transformations are calculated as follows: , In the formula, for strategy The following expectations; for The optimal state value function at time t; Define the loss function of the Actor network in the reinforcement learning algorithm. Loss function of Critic network The following formulas are shown respectively: , , In the formula, express From the experience replay pool ; This represents the random policy distribution output by the Actor network in the SAC algorithm. To reinforce the parameters of the Actor network during learning; To enhance the parameters of the Critic network during learning; Action value function The estimated value.

[0017] The second technical solution of this invention is achieved through the following method: an adaptive control system for a double-swing bridge crane based on deep reinforcement learning, using an adaptive control method for a double-swing bridge crane based on deep reinforcement learning, comprising: The model building unit is used to model the friction of the trolley, the damping of the hook and the cargo during the movement of the double pendulum bridge crane, and to build a mathematical model of the double pendulum bridge crane. The target setting unit sets the target position data of the trolley in the double pendulum bridge crane, and inputs it into the PID controller and reinforcement learning algorithm respectively. The crane drive unit, with its PID controller, receives target position data and outputs the trolley's driving force to the double-swing bridge crane's mathematical model and reinforcement learning algorithm, driving the double-swing bridge crane to move. The data optimization unit collects data on the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo during the movement of the double pendulum bridge crane, and outputs it to the reinforcement learning algorithm. Using this data, combined with the target position data of the trolley and the driving force of the trolley, the reinforcement learning algorithm is trained, the parameters of the PID controller are optimized and updated, and the optimized parameters are output to the PID controller. The error calculation unit outputs the displacement data of the trolley during the movement of the double pendulum bridge crane to the PID controller. After receiving the optimized parameters, the PID controller compares the received displacement data of the trolley with the target position data of the trolley and calculates the error between the displacement data and the target position data of the trolley. The error judgment unit determines whether the error is equal to zero. If the error is not zero, the PID controller continues to output the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the double pendulum bridge crane to move. The above process is repeated until the error between the displacement data of the trolley and the target position data of the trolley is zero, which means that the trolley has reached the target position.

[0018] This invention considers the potential double-swing effect of cranes during engineering operations and fully models the friction of the trolley, the damping of the hook and the cargo, ensuring the effectiveness and engineering applicability of the controller design. Furthermore, the reinforcement learning adaptive method of this invention is unique, differing from reinforcement learning algorithms directly used in drive systems. This research fully considers the full-state information of the double-swing crane, constructs a performance index function that balances trolley motion, system sway, and suppresses system fluctuations, and outputs three actions to optimize the system. Compared to traditional PID control methods, this adaptive control method can achieve rapid positioning and effective anti-swing while improving the system's adaptability to parameter perturbations and external disturbances. This invention addresses practical engineering problems, and its results will provide a reference for high-performance anti-swing control of double-swing bridge cranes, possessing significant application value. Attached Figure Description

[0019] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0020] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention.

[0021] Figure 2 This is the overall control flowchart of Embodiment 1 of the present invention.

[0022] Figure 3 This is a structural diagram of the double pendulum bridge crane system in Embodiment 1 of the present invention.

[0023] Figure 4 This is a flowchart of the reinforcement learning algorithm in Embodiment 1 of the present invention.

[0024] Figure 5 This is a system block diagram of Embodiment 2 of the present invention.

[0025] Figure 6 This is a comparison diagram of the effects under parameter perturbation in Embodiment 3 of the present invention.

[0026] Figure 7 This is a comparison chart of PID optimization parameters under parameter perturbation in Embodiment 3 of the present invention.

[0027] Figure 8 This is a comparison diagram of the effects under external interference in Embodiment 3 of the present invention.

[0028] Figure 9 This is a comparison chart of PID optimization parameters under external disturbances in Embodiment 3 of the present invention. Detailed Implementation

[0029] The present invention is not limited to the following embodiments, and specific implementation methods can be determined according to the technical solutions and actual conditions of the present invention.

[0030] Example 1: As Figure 1-2 As shown in the figure, this invention discloses an adaptive control method for a double-swing bridge crane based on deep reinforcement learning, including the following steps: Step S101: Model the friction of the trolley, the damping of the hook and the cargo during the movement of the double pendulum bridge crane, and establish a mathematical model of the double pendulum bridge crane. Step S102: Set the target position data of the trolley in the double pendulum bridge crane and input it into the PID controller and reinforcement learning algorithm respectively; In step S103, the PID controller receives the target position data and outputs the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the double pendulum bridge crane to move. Step S104: Collect data on the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo during the movement of the double pendulum bridge crane, and output them to the reinforcement learning algorithm. Use this data, combined with the target position data of the trolley and the driving force of the trolley, to train the reinforcement learning algorithm, optimize and update the parameters of the PID controller, and output the optimized parameters to the PID controller. Step S105: The displacement data of the trolley during the movement of the double pendulum bridge crane is output to the PID controller. After receiving the optimized parameters, the PID controller compares the received displacement data of the trolley with the target position data of the trolley and calculates the error between the displacement data of the trolley and the target position data of the trolley. Step S106: Determine if the error is zero. If no, the PID controller continues to output the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the double pendulum bridge crane to move. Repeat the above process until the error between the displacement data of the trolley and the target position data of the trolley is zero, which means that the trolley has reached the target position.

[0031] In step S101 above, a mathematical model of the double-swing bridge crane is established, modeling the friction of the trolley, the damping of the hook and the cargo during the movement of the crane. This model includes: By applying the Euler-Lagrange method, the dynamic equations of the double-pendulum bridge crane are obtained as follows: , In the formula, For the mass of the car, For the mass of the hook, For the quality of the goods, This refers to the length of the rope connecting the trolley and the hook. The length of the rope connecting the hook and the cargo. For the displacement of the car, The swing angle of the hook, The swing angle of the goods. , and For the system's damping parameters, Represents gravitational acceleration. This indicates the driving force of the car.

[0032] In the dynamic equations of the above-mentioned double pendulum bridge crane, the double pendulum effect that may occur during the operation of the crane is taken into account, and the damping of the trolley friction, hook and cargo is fully modeled. In the dynamic equations, equation (1) takes into account the trolley displacement, i.e. the damping of the trolley friction, equation (2) takes into account the damping of the hook, and equation (3) takes into account the damping of the cargo. Thus, by fully modeling the damping of the trolley friction, hook and cargo, the effectiveness and engineering applicability of the control method design of the present invention are guaranteed.

[0033] The double-swing bridge crane system includes a trolley, hook, and cargo, forming a two-stage swing structure, such as... Figure 3 As shown, during the trolley's movement, the hook and cargo inevitably oscillate. Therefore, by employing the Euler-Lagrange method, the dynamic equations of the bridge crane are derived. These equations reveal that the double-pendulum system exhibits underactuated characteristics, containing a single control input and a multi-degree-of-freedom coupled output. The trolley's movement inevitably leads to the back-and-forth oscillation of the hook and cargo. However, relying solely on air damping, this oscillation would require a very long time to eliminate. Therefore, this invention ensures both rapid and precise trolley positioning while minimizing the oscillation of the hook and cargo.

[0034] In step S102 above, the PID controller receives the target position data and outputs the driving force of the trolley to the mathematical model of the double pendulum bridge crane and the reinforcement learning SAC algorithm, including: The driving force of the car at the current moment is calculated as follows: , In the formula, Let be the driving force of the car at time t; For small car The driving force of every moment; The change in driving force between two time points. The calculation method is as follows: , In the formula, This represents the parameters of the PID controller at time t. This represents the deviation between the target position and the actual displacement of the t vehicle at time t. ,in, This indicates the target position of the vehicle. This represents the displacement of the t vehicle at time t.

[0035] The parameters of the PID controller at the next moment are as follows: , and , , and They are respectively , and The change in quantity.

[0036] In step S104 above, such as Figure 4 As shown, data on the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo during the movement of the double-pendulum bridge crane are collected and output to the reinforcement learning algorithm. This data, combined with the trolley's target position data and driving force, is used to train the reinforcement learning algorithm, optimizing and updating the parameters of the PID controller, including: Step 1: Set the current status of the double swing bridge crane system. The input is fed into the Actor network of the reinforcement learning algorithm, and the Actor network determines the input based on the current state. Calculate the actual action ; Step 2: The Actor network performs the actual actions. It interacts with the double-swing bridge crane to obtain a reward function. New Status of Double Swing Bridge Crane Systems ; Step 3: Put The data is placed into the experience replay pool. When the experience data in the experience replay pool reaches a certain amount, the initial data is deleted, and the experience data is fixed and saved to the replay pool. Step 4: Use importance sampling weighting techniques to extract a batch of empirical data from the experience replay pool, i.e., the historical batch. Calculate the values ​​of the two Critic networks and compare them, select the smaller value and update the Critic network; Step 5: Update the target Critic network and Actor network, then return to Step 1 and repeat this process until the set number of training iterations or performance metric values ​​are reached.

[0037] The displacement data of the trolley can be measured by an encoder (external sensor), and the swing angle of the hook and the swing angle of the cargo can be measured by a suspended camera or encoder.

[0038] In step 1 above, the current state of the double-swing bridge crane system is... The input is fed into the Actor network of the reinforcement learning algorithm, and the Actor network determines the input based on the current state. Calculate the actual action ,include: Based on the agent's current state Computing the optimal strategy of reinforcement learning algorithms and set The optimal action output value is calculated as follows: , In the formula, As a strategy, To determine the optimal value of the strategy, a strategy is selected based on the formula argmax. The policy value that makes the function in this optimal is the optimal policy. ; This indicates that the function The largest value; Represents the total state variables of the double-swing bridge crane system. ; This represents the actual action output by the reinforcement learning algorithm. , Let t be the strategy of the action network in the reinforcement learning algorithm; To reinforce the learning algorithm policy in the state Entropy below; These are the weighting coefficients.

[0039] Specifically, the total state variable of the double-swing bridge crane system is the displacement of the trolley. The swing angle of the hook The swing angle of the goods The target location of the car The deviation between the target position and the actual displacement of the trolley The driving force of the car .

[0040] in, To reinforce the learning algorithm policy in the state Lowering the entropy helps improve the training efficiency of reinforcement learning algorithm networks.

[0041] In step 2 above, the Actor network performs the actual actions. It interacts with the double-swing bridge crane to obtain a reward function. New Status of Double Swing Bridge Crane Systems ,include: Define a performance metric function, i.e., a reward function. The calculation method is as follows: , In the formula, This indicates the deviation between the target position and the actual displacement of the trolley; The change in driving force between two time points; , and They are respectively , and The change in quantity.

[0042] In this invention, by fully considering the full-state information of the double-swing bridge crane, a performance index function is constructed that can balance trolley motion, system oscillation, and suppress system fluctuations. This function is used by reinforcement learning algorithms to evaluate and adjust the state of the double-swing bridge crane system and outputs three actions ( and This invention is used to optimize the double pendulum bridge crane system. As a result, this invention differs from reinforcement learning algorithms that are directly used to drive the double pendulum bridge crane system. The method of this invention is unique, and through reinforcement learning algorithms, this invention achieves dynamic optimization of PID controller parameters, further improving the dynamic performance and robustness of the double pendulum bridge crane system.

[0043] In step 4 above, an important sampling weighting technique is used to extract a batch of empirical data from the experience replay pool, i.e., the historical batch. Calculate the values ​​of the two Critic networks and compare them, selecting the smaller value and updating the Critic network; including: Calculate the optimal state-value function in reinforcement learning algorithms and action value function The calculation method is as follows: , , In the formula, This is the discount factor. Indicates the strategy The expected outcome. Here, policy entropy is introduced into the state-value function. It can encourage it to maintain a certain degree of randomness, thereby avoiding getting trapped in local optima, while enhancing its robustness to external disturbances; Calculate the value of the Critic network, i.e., the action-value function. Further transformations are calculated as follows: , In the formula, for strategy The expected value is lowered, which effectively reduces the variance of the solution, makes the network update more stable, and avoids the oscillation phenomenon caused by a strong double pendulum effect, thus improving the adaptability of the control system. Furthermore, for The optimal state value function at time t; Define the loss function of the Actor network in the reinforcement learning algorithm. Loss function of Critic network The following formulas are shown respectively: , , In the formula, express From the experience replay pool ; This represents the random policy distribution output by the Actor network in the SAC algorithm. To reinforce the parameters of the Actor network during learning; To enhance the parameters of the Critic network during learning; Action value function The estimated value is obtained. Furthermore, the multiple Q-value minimization method in the above formula effectively suppresses the Q-value overestimation problem commonly found in deep reinforcement learning, increasing the accuracy and stability of policy updates.

[0044] In summary, the embodiments of this invention, by considering the potential double-swing effect of the crane during engineering operations and fully modeling the friction of the trolley, the damping of the hook and the cargo, ensure the effectiveness and engineering applicability of the controller design. Furthermore, the reinforcement learning adaptive method of this invention is unique, differing from reinforcement learning algorithms directly used in driving systems. This research fully considers the full-state information of the double-swing crane, constructs a performance index function that can balance trolley motion, system sway, and suppress system fluctuations, and outputs three actions to optimize the system. Compared to traditional PID control methods, the adaptive control method of this invention can achieve rapid positioning and effective anti-swing while improving the system's adaptability to parameter perturbations and external disturbances. This invention addresses practical engineering problems, and its results will provide a reference for high-performance anti-swing control of double-swing bridge cranes, possessing significant application value.

[0045] Example 2: As Figure 5 As shown, this invention discloses an adaptive control system for a double-swing bridge crane based on deep reinforcement learning. The method utilizes a deep reinforcement learning-based adaptive control approach for the double-swing bridge crane, comprising: The model building unit is used to model the friction of the trolley, the damping of the hook and the cargo during the movement of the double pendulum bridge crane, and to build a mathematical model of the double pendulum bridge crane. The target setting unit sets the target position data of the trolley in the double pendulum bridge crane, and inputs it into the PID controller and reinforcement learning algorithm respectively. The crane drive unit, with its PID controller, receives target position data and outputs the trolley's driving force to the double-swing bridge crane's mathematical model and reinforcement learning algorithm, driving the double-swing bridge crane to move. The data optimization unit collects data on the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo during the movement of the double pendulum bridge crane, and outputs it to the reinforcement learning algorithm. Using this data, combined with the target position data of the trolley and the driving force of the trolley, the reinforcement learning algorithm is trained, the parameters of the PID controller are optimized and updated, and the optimized parameters are output to the PID controller. The error calculation unit outputs the displacement data of the trolley during the movement of the double pendulum bridge crane to the PID controller. After receiving the optimized parameters, the PID controller compares the received displacement data of the trolley with the target position data of the trolley and calculates the error between the displacement data and the target position data of the trolley. The error judgment unit determines whether the error is equal to zero. If the error is not zero, the PID controller continues to output the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the double pendulum bridge crane to move. The above process is repeated until the error between the displacement data of the trolley and the target position data of the trolley is zero, which means that the trolley has reached the target position.

[0046] Example 3: In this example, the mechanical parameters of the bridge crane are selected as follows: , , , , , and Furthermore, the deep reinforcement learning algorithm parameters were set as follows: learning rate of 0.0001, experience replay buffer capacity of 20000, initial entropy weight coefficient of 0.01, 500 training rounds, 200 training iterations per round, 40s simulation time per simulation, and 0.01s simulation step size. The number of nodes in the input, hidden, and output layers of the three Actor networks in the deep reinforcement learning algorithm were 6, 256, and 1, respectively; the number of nodes in the input, hidden layer 1, hidden layer 2, and output layers of the two Critic networks and two target Critic networks were 9, 256, 256, and 1, respectively. This study compared this adaptive PID control with traditional PID control under two scenarios: parameter perturbation and external disturbance.

[0047] In the parameter perturbation experiment, the hook mass was set to 20kg, the load mass to 40kg, the rope length between the trolley and the hook to 4m, and the rope length between the hook and the load to 2m. The system dynamic response results are as follows: Figure 6As shown in the figure, the dynamic performance of the crane system differs significantly under the two control methods. The trolley's settling time is approximately 15 seconds and 10 seconds under these two control algorithms, respectively; the maximum swing angle of the hook is approximately 9° and 6°, respectively; and the maximum swing angle of the load is approximately 11° and 9°, respectively. Due to the change in rope length, the swing angle is reduced under PID algorithm control. Even so, the PID algorithm after parameter optimization, i.e., reinforcement learning adaptive PID, has advantages in swing angle suppression. Figure 7 As shown, the optimized PID parameters stabilize at , , Compared with fixed parameters, its proportional and differential parameters change significantly, especially the differential parameters.

[0048] Consider the external disturbances such as gusts experienced by the crane system in actual engineering. Assume that at the 20th second, a disturbance of 20N with a duration of 2 seconds is introduced; starting from the 30th second, a continuous disturbance of 10sin(t)N is introduced. The system dynamic response is as follows: Figure 8 As shown in the figure, compared to PID control, the adaptive algorithm results in less trolley fluctuation and significantly reduced swaying of the hook and cargo. This is due to the adaptive algorithm's real-time and rapid optimization of PID parameters, such as... Figure 9 As shown above, the verification of the two uncertainties demonstrates that the crane system under reinforcement learning adaptive PID algorithm control has strong disturbance suppression capability, ensuring good dynamic and steady-state performance.

Claims

1. An adaptive control method for a double-pendulum bridge crane based on deep reinforcement learning, characterized in that, Includes the following steps: A mathematical model of the double pendulum bridge crane is established by modeling the friction of the trolley, the damping of the hook and the cargo during the movement of the double pendulum bridge crane. Set the target position data of the trolley in the double pendulum bridge crane and input it into the PID controller and reinforcement learning algorithm respectively; The PID controller receives the target position data and outputs the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the movement of the double pendulum bridge crane. Data on the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo during the movement of the double pendulum bridge crane are collected and output to the reinforcement learning algorithm. The reinforcement learning algorithm is trained using this data, combined with the target position data of the trolley and the driving force of the trolley, to optimize and update the parameters of the PID controller. The optimized parameters are then output to the PID controller. The displacement data of the trolley during the movement of the double pendulum bridge crane is output to the PID controller. After receiving the optimized parameters, the PID controller compares the received displacement data of the trolley with the target position data of the trolley and calculates the error between the displacement data and the target position data of the trolley. The system determines whether the error is zero. If no, the PID controller continues to output the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the double pendulum bridge crane to move. The above process is repeated until the error between the displacement data of the trolley and the target position data of the trolley is zero, which means that the trolley has reached the target position.

2. The adaptive control method for a double-pendulum bridge crane based on deep reinforcement learning according to claim 1, characterized in that, The process involves modeling the friction of the trolley, the damping of the hook and the cargo during the movement of a double-swing bridge crane, establishing a mathematical model of the double-swing bridge crane, including: By applying the Euler-Lagrange method, the dynamic equations of the double-pendulum bridge crane are obtained as follows: , In the formula, For the mass of the car, For the mass of the hook, For the quality of the goods, This refers to the length of the rope connecting the trolley and the hook. The length of the rope connecting the hook and the cargo. For the displacement of the car, The swing angle of the hook, The swing angle of the goods. , and For the system's damping parameters, Represents gravitational acceleration. This indicates the driving force of the car.

3. The adaptive control method for a double-pendulum bridge crane based on deep reinforcement learning according to claim 1, characterized in that, The PID controller receives target position data and outputs the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double-swing bridge crane, including: The driving force of the car at the current moment is calculated as follows: , In the formula, Let be the driving force of the car at time t; For small car The driving force of every moment; The change in driving force between two time points. The calculation method is as follows: , In the formula, This represents the parameters of the PID controller at time t. This represents the deviation between the target position and the actual displacement of the t vehicle at time t. ,in, This indicates the target position of the vehicle. This represents the displacement of the t vehicle at time t.

4. The adaptive control method for a double-pendulum bridge crane based on deep reinforcement learning according to claim 1, characterized in that, The data collected during the movement of the double-swing bridge crane include the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo. This data is then output to a reinforcement learning algorithm. Using this data, combined with the trolley's target position data and driving force, the reinforcement learning algorithm is trained, and the parameters of the PID controller are optimized and updated. This includes: Step 1: Set the current status of the double swing bridge crane system. The input is fed into the Actor network of the reinforcement learning algorithm, and the Actor network determines the input based on the current state. Calculate the actual action ; Step 2: The Actor network performs the actual actions. It interacts with the double-swing bridge crane to obtain a reward function. New Status of Double Swing Bridge Crane Systems ; Step 3: Put The data is placed into the experience replay pool. When the experience data in the experience replay pool reaches a certain amount, the initial data is deleted, and the experience data is fixed and saved to the replay pool. Step 4: Use importance sampling weighting techniques to extract a batch of empirical data from the experience replay pool, i.e., the historical batch. Calculate the values ​​of the two Critic networks and compare them, select the smaller value and update the Critic network; Step 5: Update the target Critic network and Actor network, then return to Step 1 and repeat this process until the set number of training iterations or performance metric values ​​are reached.

5. The adaptive control method for a double-pendulum bridge crane based on deep reinforcement learning according to claim 4, characterized in that, The current state of the double swing bridge crane system The input is fed into the Actor network of the reinforcement learning algorithm, and the Actor network determines the input based on the current state. Calculate the actual action ,include: Based on the agent's current state Computing the optimal strategy of reinforcement learning algorithms and set The optimal action output value is calculated as follows: , In the formula, As a strategy, This represents the optimal value of the strategy. This indicates that the function The largest value; Represents the total state variables of the double-swing bridge crane system. ; This represents the actual action output by the reinforcement learning algorithm. , Let t be the strategy of the action network in the reinforcement learning algorithm; To reinforce the learning algorithm policy in the state Entropy below; These are the weighting coefficients.

6. The adaptive control method for a double-pendulum bridge crane based on deep reinforcement learning according to claim 4, characterized in that, The Actor network performs actual actions. It interacts with the double-swing bridge crane to obtain a reward function. New Status of Double Swing Bridge Crane Systems ,include: Define a performance metric function, i.e., a reward function. The calculation method is as follows: , In the formula, This indicates the deviation between the target position and the actual displacement of the trolley; The change in driving force between two time points; , and They are respectively , and The change in quantity.

7. The adaptive control method for a double-pendulum bridge crane based on deep reinforcement learning according to claim 4, characterized in that, The method utilizes importance sampling weighting to extract a batch of empirical data from the experience replay pool, i.e., historical batches. Calculate the values ​​of the two Critic networks and compare them, selecting the smaller value and updating the Critic network; including: Calculate the optimal state-value function in reinforcement learning algorithms and action value function The calculation method is as follows: , , In the formula, This is the discount factor. Indicates the strategy The following expectations; Calculate the value of the Critic network, i.e., the action-value function. Further transformations are calculated as follows: , In the formula, for strategy The following expectations; for The optimal state value function at time t; Define the loss function of the Actor network in the reinforcement learning algorithm. Loss function of Critic network The following formulas are shown respectively: , , In the formula, express From the experience replay pool ; This represents the random policy distribution output by the Actor network in the SAC algorithm. To reinforce the parameters of the Actor network during learning; To enhance the parameters of the Critic network during learning; Action value function The estimated value.

8. An adaptive control system for a double-swing bridge crane based on deep reinforcement learning, using the adaptive control method for a double-swing bridge crane based on deep reinforcement learning as described in any one of claims 1 to 7, characterized in that, include: The model building unit is used to model the friction of the trolley, the damping of the hook and the cargo during the movement of the double pendulum bridge crane, and to build a mathematical model of the double pendulum bridge crane. The target setting unit sets the target position data of the trolley in the double pendulum bridge crane, and inputs it into the PID controller and reinforcement learning algorithm respectively. The crane drive unit, with its PID controller, receives target position data and outputs the trolley's driving force to the double-swing bridge crane's mathematical model and reinforcement learning algorithm, driving the double-swing bridge crane to move. The data optimization unit collects data on the displacement of the trolley, the swing angle of the hook, and the swing angle of the cargo during the movement of the double pendulum bridge crane, and outputs it to the reinforcement learning algorithm. Using this data, combined with the target position data of the trolley and the driving force of the trolley, the reinforcement learning algorithm is trained, the parameters of the PID controller are optimized and updated, and the optimized parameters are output to the PID controller. The error calculation unit outputs the displacement data of the trolley during the movement of the double pendulum bridge crane to the PID controller. After receiving the optimized parameters, the PID controller compares the received displacement data of the trolley with the target position data of the trolley and calculates the error between the displacement data and the target position data of the trolley. The error judgment unit determines whether the error is equal to zero. If the error is not zero, the PID controller continues to output the driving force of the trolley to the mathematical model and reinforcement learning algorithm of the double pendulum bridge crane to drive the double pendulum bridge crane to move. The above process is repeated until the error between the displacement data of the trolley and the target position data of the trolley is zero, which means that the trolley has reached the target position.

Citation Information

Patent Citations

  • Fuzzy PID control system and method for lifting appliance of active anti-swing crane

    CN113336093A

  • Bridge crane sling anti-swing method based on trajectory planning

    CN114955856A