Bridge portal crane anti-sway control method and device based on deep reinforcement learning

CN117466145BActive Publication Date: 2026-08-11WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但深度强化学习在桥门式起重机防摇控制方面的研究尚不充分,如何将深度强化学习应用于桥式桥门式起重机吊具的防摇控制,成为亟需解决的问题

Benefits of technology

[0036] Compared with the prior art, the beneficial effects of the present invention include: firstly, constructing a virtual platform for anti-sway control of a gantry crane, and using an input shaping algorithm to determine the initial strategy of the virtual platform; then, using a deep deterministic policy gradient algorithm to optimize the initial strategy of the virtual platform to obtain the final strategy of the virtual platform; finally, using a double-Q network to transfer the final strategy of the virtual platform to the real platform for anti-sway control of the gantry crane to obtain the anti-sway control strategy of the real platform for anti-sway control of the gantry crane. This realizes the application of deep reinforcement learning algorithms in anti-sway control of gantry cranes and improves the performance of anti-sway control of gantry cranes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117466145B_ABST
    Figure CN117466145B_ABST
Patent Text Reader

Abstract

This invention relates to a method and apparatus for anti-sway control of gantry cranes based on deep reinforcement learning, comprising: constructing a virtual platform for anti-sway control of the gantry crane; determining an initial strategy for the virtual platform based on an input shaping algorithm; determining a final strategy for the virtual platform based on the initial strategy and a deep deterministic policy gradient algorithm; and transferring the final strategy to a real platform for anti-sway control of the gantry crane based on a dual-Q network to determine the anti-sway control strategy of the real platform. This invention realizes the application of deep reinforcement learning algorithms in the anti-sway control of gantry cranes, improving the performance of anti-sway control for gantry cranes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of lifting and transportation technology, specifically to a method and device for anti-sway control of bridge and gantry cranes based on deep reinforcement learning. Background Technology

[0002] Gantry cranes are lifting and transporting equipment whose lifting devices are arranged on horizontal beams spanning workshops and storage yards. They are widely used in industrial sites such as workshops, ports, and warehouses. Based on the operating scenario, they can be divided into industrial gantry cranes, rail-mounted gantry cranes, railway gantry cranes, and container gantry cranes. The trolley traveling mechanism and the lifting mechanism of a gantry crane are connected by flexible steel wire ropes. When the trolley and hoisting mechanisms use variable speed drives, due to inertial forces and external interference factors such as wind, the spreader will experience an approximate pendulum motion. This swaying will severely affect the positioning accuracy of the spreader, increase the difficulty of stacking goods, and reduce the loading and unloading efficiency of the gantry crane. Furthermore, excessive swaying may lead to dangerous accidents. Therefore, to mitigate this situation, gantry cranes need to be equipped with anti-sway devices. Currently, commonly used anti-sway methods mainly include manual anti-sway, mechanical anti-sway, and electronic anti-sway.

[0003] In recent years, deep reinforcement learning has gradually attracted attention as a method suitable for handling complex nonlinear systems. Deep reinforcement learning can learn optimal control strategies based on the environment and external rewards, and can adaptively handle unknown parameters and dynamic influences. However, research on the anti-sway control of bridge and gantry cranes is still insufficient. How to apply deep reinforcement learning to the anti-sway control of bridge and gantry crane spreaders has become an urgent problem to be solved. Summary of the Invention

[0004] In view of this, it is necessary to provide a method and device for anti-sway control of bridge and gantry cranes based on deep reinforcement learning, so as to solve the technical problem that it is difficult to apply deep reinforcement learning to the anti-sway control of bridge and gantry crane spreaders in the current technology.

[0005] To achieve the above objectives, this invention provides a method for anti-sway control of bridge gantry cranes based on deep reinforcement learning, comprising:

[0006] A virtual platform for anti-sway control of a bridge crane is constructed, and the initial strategy of the virtual platform is determined based on an input shaping algorithm.

[0007] Based on the initial strategy and the deep deterministic strategy gradient algorithm, the final strategy of the anti-sway control virtual platform for the bridge crane is determined.

[0008] Based on the dual-Q network, the final strategy is transferred to the real platform for anti-sway control of bridge and gantry cranes to determine the anti-sway control strategy of the real platform for anti-sway control of bridge and gantry cranes.

[0009] Furthermore, the initial strategy for determining the anti-sway control virtual platform for the bridge crane based on the input shaping algorithm includes:

[0010] The initial strategy is determined based on the ZV input shaping algorithm, the natural frequency of the spreader oscillation in the anti-sway control virtual platform of the gantry crane, and the damping ratio of the system in the anti-sway control virtual platform of the gantry crane.

[0011] Further, the initial strategy is determined based on the ZV input shaping algorithm, the natural frequency of the spreader oscillation in the anti-sway control virtual platform of the gantry crane, and the damping ratio of the system in the anti-sway control virtual platform of the gantry crane, including:

[0012] The initial strategy is determined based on the following formula:

[0013]

[0014] Where A1 represents the amplitude of the first pulse initiated by the anti-sway control virtual platform for the gantry crane, A2 represents the amplitude of the second pulse initiated by the anti-sway control virtual platform for the gantry crane, t1 represents the triggering time of the first pulse, t2 represents the triggering time of the second pulse, and ω n The natural frequency of the sway of the spreader in the anti-sway control virtual platform of the gantry crane is represented by ξ, the damping ratio of the system in the anti-sway control virtual platform of the gantry crane is represented by ξ, and K is a proportional parameter. The first pulse and the second pulse have the same duration. The first pulse and the second pulse are used to drive the trolley in the anti-sway control virtual platform of the gantry crane.

[0015] Furthermore, the determination of the final strategy for the anti-sway control virtual platform of the bridge crane based on the initial strategy and the deep deterministic strategy gradient algorithm includes:

[0016] Based on the initial strategy, the state and acceleration of the system in the anti-sway control virtual platform for the bridge crane at any given time are determined, as well as the state of the system in the anti-sway control virtual platform for the bridge crane at the next time after that time. Based on the state and acceleration of the system in the anti-sway control virtual platform for the bridge crane at any given time, the reward of the system in the anti-sway control virtual platform for the bridge crane at any given time is determined.

[0017] Based on the state, acceleration, and reward of the system in the anti-sway control virtual platform for bridge and gantry cranes at any given moment, and the state of the system in the anti-sway control virtual platform for bridge and gantry cranes at the next moment after that moment, an offline experience base is constructed.

[0018] Using the offline experience base as an experience replay pool, the final strategy of the anti-sway control virtual platform for the bridge crane is determined based on the deep deterministic policy gradient algorithm.

[0019] Further, determining the reward of the system in the anti-sway control virtual platform for the gantry crane at any given time, based on the system's state and acceleration at any given moment, includes:

[0020] When the trolley in the anti-sway control virtual platform for the gantry crane is in operation, the reward of the system in the anti-sway control virtual platform for the gantry crane at any given moment is determined based on the following formula:

[0021]

[0022] If the trolley in the anti-sway control virtual platform for the gantry crane reaches the destination, the reward of the system in the anti-sway control virtual platform for the gantry crane at any given moment is determined based on the following formula:

[0023] R(s t ,a t ) = 10*(5-n)

[0024] Wherein, R(s) t ,a t ) represents the reward of the system in the anti-sway control virtual platform for the bridge crane at any given time, s t This indicates the state of the system in the anti-sway control virtual platform for the bridge crane at any given time, a t The acceleration of the system in the anti-sway control virtual platform of the bridge crane at any given moment is represented by n, where n represents the number of cycles in which the swing amplitude of the spreader is less than the preset amplitude.

[0025] Furthermore, based on the dual-Q network, the final strategy is transferred to the bridge crane anti-sway control real platform to determine the anti-sway control strategy of the bridge crane anti-sway control real platform, including:

[0026] The value function network corresponding to the final strategy is used as the source network of the double-Q network, the target network of the double-Q network is randomly initialized, and the source network is updated.

[0027] Based on the updated source network, the policy function corresponding to the final policy is updated, and the updated policy function corresponding to the final policy is used as the anti-sway control policy of the bridge crane anti-sway control real platform.

[0028] Furthermore, the system status in the virtual platform for anti-sway control of the gantry crane and the system status in the real platform for anti-sway control of the gantry crane include:

[0029] The position and speed of the trolley, as well as the swing angle and angular velocity of the spreader.

[0030] The present invention also provides an anti-sway control device for bridge cranes based on deep reinforcement learning, comprising:

[0031] A construction module is used to build a virtual platform for anti-sway control of bridge and gantry cranes, and to determine the initial strategy of the virtual platform for anti-sway control of bridge and gantry cranes based on an input shaping algorithm;

[0032] The first determining module is used to determine the final strategy of the bridge crane anti-sway control virtual platform based on the initial strategy and the deep deterministic strategy gradient algorithm.

[0033] The second determining module is used to migrate the final strategy to the bridge crane anti-sway control real platform based on the dual-Q network, and determine the anti-sway control strategy of the bridge crane anti-sway control real platform.

[0034] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the anti-sway control method for bridge cranes based on deep reinforcement learning as described above.

[0035] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the deep reinforcement learning-based anti-sway control method for bridge and gantry cranes as described above.

[0036] Compared with the prior art, the beneficial effects of the present invention include: firstly, constructing a virtual platform for anti-sway control of a gantry crane, and using an input shaping algorithm to determine the initial strategy of the virtual platform; then, using a deep deterministic policy gradient algorithm to optimize the initial strategy of the virtual platform to obtain the final strategy of the virtual platform; finally, using a double-Q network to transfer the final strategy of the virtual platform to the real platform for anti-sway control of the gantry crane to obtain the anti-sway control strategy of the real platform for anti-sway control of the gantry crane. This realizes the application of deep reinforcement learning algorithms in anti-sway control of gantry cranes and improves the performance of anti-sway control of gantry cranes. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 A flowchart illustrating an embodiment of the anti-sway control method for bridge gantry cranes based on deep reinforcement learning provided by the present invention;

[0039] Figure 2 This is a flowchart illustrating an embodiment of the anti-sway method for gantry crane spreaders at the shore container bridge provided by the present invention.

[0040] Figure 3 A schematic diagram of the structure of an embodiment of the quay crane miniature model platform provided by the present invention;

[0041] Figure 4 This is a schematic diagram of an embodiment of the input pre-shaping trolley acceleration provided by the present invention;

[0042] Figure 5 This is a schematic diagram of an embodiment of the input-shaped vehicle acceleration provided by the present invention;

[0043] Figure 6 A schematic diagram of an embodiment of the input-shaped vehicle speed provided by the present invention;

[0044] Figure 7 This is a flowchart illustrating an embodiment of the DDPG reinforcement learning algorithm provided by the present invention;

[0045] Figure 8 A flowchart illustrating an embodiment of the anti-shake algorithm migration from virtual experiments to real environments provided by the present invention;

[0046] Figure 9 A schematic diagram of an embodiment of the anti-sway control device for bridge gantry cranes based on deep reinforcement learning provided by the present invention;

[0047] Figure 10 A schematic diagram of the structure of an embodiment of the electronic device provided by the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0049] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. Furthermore, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0050] In the description of this invention, reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the described embodiments can be combined with other embodiments.

[0051] In recent years, deep reinforcement learning has gradually attracted attention as a method suitable for handling complex nonlinear systems. Deep reinforcement learning can learn optimal control strategies based on the environment and external rewards, and can adaptively handle unknown parameters and dynamic influences. Although deep reinforcement learning has been widely applied in industrial manufacturing, robot control, scheduling optimization, and game theory, research on anti-sway control of gantry cranes is still insufficient. Applying deep reinforcement learning to the anti-sway control of gantry crane spreaders holds promise for better handling of complex nonlinear systems and providing superior control performance.

[0052] Eight-rope anti-sway is a representative method of mechanical anti-sway. It uses mechanical means to dissipate the energy of load swaying during trolley operation, thereby ultimately reducing load sway. This method has advantages in stability and reliability, but it also brings some problems, such as increased overall mass of the gantry crane, high energy consumption, difficult maintenance, and high hardware costs.

[0053] With the continuous improvement of automation in port gantry cranes, electronic anti-sway technology has been widely applied and has become the main control method for anti-sway systems of port gantry crane spreaders. This anti-sway method is based on control theory. By analyzing the relationship between various states of the gantry crane, it calculates the input signal that can ensure accurate positioning and anti-sway of the system. Theoretically, electronic anti-sway is more effective and less expensive. Currently, research on electronic anti-sway is mainly divided into open-loop control and closed-loop control. Open-loop electronic anti-sway is mainly achieved through theoretical methods such as given speed curves and optimal control. It does not rely on angle feedback, so it is less expensive and easier to implement. However, it requires accurate modeling, mainly using methods such as input shaping. Closed-loop electronic anti-sway measures the changes in the controlled variable by installing sensors and feeding this information back. Based on the feedback results, the system output is adjusted in real time to achieve more accurate control. However, conventional closed-loop control requires comprehensive feedback states, and environmental influences such as wind and waves in ports make it difficult for traditional closed-loop control to achieve good results.

[0054] To overcome the problems of existing anti-sway control methods, such as difficulty in simplifying modeling, difficulty in obtaining full-state feedback, and inability to translate theoretical simulation research into practical engineering applications, this invention proposes an anti-sway control method for bridge and gantry cranes based on deep reinforcement learning to suppress the swaying of the suspended load.

[0055] The specific embodiments are described in detail below:

[0056] This invention provides a method for anti-sway control of bridge gantry cranes based on deep reinforcement learning, combined with... Figure 1 Let's take a look. Figure 1 This is a flowchart illustrating an embodiment of the anti-sway control method for bridge gantry cranes based on deep reinforcement learning provided by the present invention, including steps S101 to S103, wherein:

[0057] In step S101, a virtual platform for anti-sway control of a gantry crane is constructed, and the initial strategy of the virtual platform for anti-sway control of the gantry crane is determined based on an input shaping algorithm.

[0058] In step S102, the final strategy of the anti-sway control virtual platform for the bridge crane is determined based on the initial strategy and the deep deterministic strategy gradient algorithm.

[0059] In step S103, based on the dual-Q network, the final strategy is transferred to the bridge crane anti-sway control real platform to determine the anti-sway control strategy of the bridge crane anti-sway control real platform.

[0060] In this embodiment of the invention, a virtual platform for anti-sway control of a gantry crane is first constructed, and an input shaping algorithm is used to determine the initial strategy of the virtual platform. Then, a deep deterministic policy gradient algorithm is used to optimize the initial strategy of the virtual platform to obtain the final strategy of the virtual platform. Finally, a double-Q network is used to transfer the final strategy of the virtual platform to the real platform for anti-sway control of the gantry crane to obtain the anti-sway control strategy of the real platform for anti-sway control of the gantry crane. This realizes the application of deep reinforcement learning algorithm in anti-sway control of gantry cranes and improves the performance of anti-sway control of gantry cranes.

[0061] In a specific embodiment of the present invention, a virtual platform for anti-sway control of a gantry crane can first be constructed to simulate the usage scenario of a quay container spreader. In this scenario, during a single operation, the trolley typically accelerates first, maintains a constant speed in the middle, and then decelerates. An input shaping algorithm can be used to shape the acceleration of the trolley, and the result is used as the initial strategy of the virtual platform for anti-sway control of the gantry crane.

[0062] After obtaining the initial strategy of the anti-sway control virtual platform for the gantry crane, the Deep Deterministic Policy Gradient (DDPG) algorithm can be used to optimize the initial strategy of the anti-sway control virtual platform for the gantry crane. This allows the anti-sway control virtual platform to output more appropriate actions based on the current state of the system, and finally obtain the final strategy of the anti-sway control virtual platform for the gantry crane.

[0063] Since the strategy obtained by DDPG may be overestimated, when migrating the final strategy of the virtual platform for anti-sway control of gantry cranes to the real platform for anti-sway control of gantry cranes, a dual-Q network can be used to reduce the overestimation of the final strategy of the virtual platform for anti-sway control of gantry cranes, so as to obtain the anti-sway control strategy of the real platform for anti-sway control of gantry cranes.

[0064] In a preferred embodiment, the initial strategy for determining the anti-sway control virtual platform for the gantry crane based on the input shaping algorithm includes:

[0065] The initial strategy is determined based on the ZV input shaping algorithm, the natural frequency of the spreader oscillation in the anti-sway control virtual platform of the gantry crane, and the damping ratio of the system in the anti-sway control virtual platform of the gantry crane.

[0066] In a specific embodiment of the present invention, when using the input shaping algorithm to determine the initial strategy of the anti-sway control virtual platform for a gantry crane, the ZV input shaping algorithm can be used to determine the acceleration of the trolley during a single operation based on the natural frequency of the spreader swing and the damping ratio of the system in the anti-sway control virtual platform for the gantry crane, thereby determining the initial strategy of the anti-sway control virtual platform for the gantry crane.

[0067] In a preferred embodiment, determining the initial strategy based on the ZV input shaping algorithm, the natural frequency of the spreader oscillation in the anti-sway control virtual platform of the gantry crane, and the damping ratio of the system in the anti-sway control virtual platform of the gantry crane includes:

[0068] The initial strategy is determined based on the following formula:

[0069]

[0070] Where A1 represents the amplitude of the first pulse initiated by the anti-sway control virtual platform for the gantry crane, A2 represents the amplitude of the second pulse initiated by the anti-sway control virtual platform for the gantry crane, t1 represents the triggering time of the first pulse, t2 represents the triggering time of the second pulse, and ω n The natural frequency of the sway of the spreader in the anti-sway control virtual platform of the gantry crane is represented by ξ, the damping ratio of the system in the anti-sway control virtual platform of the gantry crane is represented by ξ, and K is a proportional parameter. The first pulse and the second pulse have the same duration. The first pulse and the second pulse are used to drive the trolley in the anti-sway control virtual platform of the gantry crane.

[0071] In a specific embodiment of the present invention, when using the ZV input shaping algorithm to determine the initial strategy of the gantry crane anti-sway control virtual platform based on the natural frequency of the spreader's oscillation and the damping ratio of the system in the gantry crane anti-sway control virtual platform, the above formula can be used. The first pulse and the second pulse can be used to drive the trolley in the gantry crane anti-sway control virtual platform, thereby providing acceleration to the trolley.

[0072] In a preferred embodiment, determining the final strategy of the bridge crane anti-sway control virtual platform based on the initial strategy and the deep deterministic strategy gradient algorithm includes:

[0073] Based on the initial strategy, the state and acceleration of the system in the anti-sway control virtual platform for the bridge crane at any given time are determined, as well as the state of the system in the anti-sway control virtual platform for the bridge crane at the next time after that time. Based on the state and acceleration of the system in the anti-sway control virtual platform for the bridge crane at any given time, the reward of the system in the anti-sway control virtual platform for the bridge crane at any given time is determined.

[0074] Based on the state, acceleration, and reward of the system in the anti-sway control virtual platform for bridge and gantry cranes at any given moment, and the state of the system in the anti-sway control virtual platform for bridge and gantry cranes at the next moment after that moment, an offline experience base is constructed.

[0075] Using the offline experience base as an experience replay pool, the final strategy of the anti-sway control virtual platform for the bridge crane is determined based on the deep deterministic policy gradient algorithm.

[0076] In a specific embodiment of the present invention, after determining the initial strategy of the anti-sway control virtual platform for the gantry crane, the state, acceleration, and reward of the system in the virtual platform at any given time, as well as the state of the system at the next time, can be determined based on the initial strategy. Taking time t as an example, the state of the system at time t is s. t The acceleration of the car is a t The state of the system at time t+1 is s. t+1 The reward received is r t It can be (s) t ,a t ,s t+1 ,r t This data is stored as a set of data in the offline experience library.

[0077] After constructing the offline experience base, it can be used as an experience replay pool for the DDPG algorithm. The DDPG algorithm is then used to optimize the initial strategy of the anti-sway control virtual platform for gantry cranes, resulting in the final strategy. During the optimization process, if the experience replay pool becomes full, data in the offline experience base can be deleted first, followed by data with earlier iteration times.

[0078] In a preferred embodiment, determining the reward of the system in the anti-sway control virtual platform for the gantry crane at any given time, based on the system's state and acceleration at any given moment, includes:

[0079] When the trolley in the anti-sway control virtual platform for the gantry crane is in operation, the reward of the system in the anti-sway control virtual platform for the gantry crane at any given moment is determined based on the following formula:

[0080]

[0081] If the trolley in the anti-sway control virtual platform for the gantry crane reaches the destination, the reward of the system in the anti-sway control virtual platform for the gantry crane at any given moment is determined based on the following formula:

[0082] R(s t ,a t ) = 10*(5-n)

[0083] Wherein, R(s) t ,a t ) represents the reward of the system in the anti-sway control virtual platform for the bridge crane at any given time, s t This indicates the state of the system in the anti-sway control virtual platform for the bridge crane at any given time, a t The acceleration of the system in the anti-sway control virtual platform of the bridge crane at any given moment is represented by n, where n represents the number of cycles in which the swing amplitude of the spreader is less than the preset amplitude.

[0084] In a specific embodiment of the present invention, the reward of the system in the anti-sway control virtual platform for bridge and gantry cranes at any given time can be determined according to the above formula.

[0085] In a preferred embodiment, the step of migrating the final strategy to the bridge crane anti-sway control real platform based on the dual-Q network, and determining the anti-sway control strategy of the bridge crane anti-sway control real platform, includes:

[0086] The value function network corresponding to the final strategy is used as the source network of the double-Q network, the target network of the double-Q network is randomly initialized, and the source network is updated.

[0087] Based on the updated source network, the policy function corresponding to the final policy is updated, and the updated policy function corresponding to the final policy is used as the anti-sway control policy of the bridge crane anti-sway control real platform.

[0088] In a specific embodiment of the present invention, when using a dual-Q network to migrate the final strategy of the virtual platform for anti-sway control of a gantry crane to the real platform for anti-sway control of a gantry crane, the value function network corresponding to the final strategy of the virtual platform for anti-sway control of the gantry crane can be used as the source network of the dual-Q network, and the target network of the dual-Q network can be randomly initialized. Then, the source network of the dual-Q network is updated, and the strategy function corresponding to the final strategy of the virtual platform for anti-sway control of the gantry crane is updated according to the updated source network. The updated strategy function is then used as the anti-sway control strategy of the real platform for anti-sway control of the gantry crane.

[0089] In a preferred embodiment, the system states in the virtual platform for anti-sway control of the gantry crane and the system states in the real platform for anti-sway control of the gantry crane include:

[0090] The position and speed of the trolley, as well as the swing angle and angular velocity of the spreader.

[0091] In a specific embodiment of the present invention, during the execution of the deep reinforcement learning algorithm, the state of the system in the virtual platform for anti-sway control of the gantry crane and the state of the system in the real platform for anti-sway control of the gantry crane may include the position and speed of the trolley, and the swing angle and angular velocity of the spreader. Correspondingly, the actions of the system may include the acceleration of the trolley.

[0092] The following specific application scenario will better illustrate the technical solution of the present invention:

[0093] Combination Figure 2 Let's take a look. Figure 2 This is a flowchart illustrating an embodiment of the anti-sway method for gantry crane spreaders provided by the present invention. The method includes three parts: prior policy learning based on input shaping, training of the spreader anti-sway algorithm in a virtual environment, and transfer of the spreader anti-sway algorithm to a real environment.

[0094] Combination Figure 3 Let's take a look. Figure 3 This is a structural schematic diagram of an embodiment of the quay crane miniature model platform provided by the present invention. The miniature experimental device for the quay crane container gantry crane includes hardware equipment such as a hoisting mechanism, a trolley mechanism, an electrical cabinet, an automated guided vehicle (AGV) trolley, containers and spreaders, and a computer.

[0095] Prior strategy learning based on input shaping for anti-sway: The idea behind input shaping is to input the original input signal in n steps (n≥2). By controlling the time interval, the vibrations generated by each signal can be canceled out through linear superposition. In essence, input shaping is calculating the amplitude and hysteresis time of each pulse signal. In existing applications, the speed curve of the trolley during a single operation is usually a trapezoidal curve with initial acceleration, then uniform speed, and finally deceleration. This invention assumes a constant rope length and uses the traditional ZV shaping input method to shape the trolley's acceleration. The constraints are as follows:

[0096]

[0097] Solving for the given information, we get:

[0098]

[0099] Where A1 represents the amplitude of the first pulse initiated by the quay container spreader anti-sway virtual test platform (i.e., the gantry crane anti-sway control virtual platform), A2 represents the amplitude of the second pulse initiated by the quay container spreader anti-sway virtual test platform, t1 represents the trigger time of the first pulse, t2 represents the trigger time of the second pulse, and ω n Let ξ represent the natural frequency of the swaying spreader, ξ represent the damping ratio of the system, and K be the proportional parameter.

[0100] Combination Figure 4 and Figure 5 Let's take a look. Figure 4 This is a schematic diagram of an embodiment of the input pre-shaping trolley acceleration provided by the present invention. Figure 5 This is a schematic diagram of an embodiment of the acceleration of a vehicle after input shaping provided by the present invention. Input shaping decomposes the acceleration process of the vehicle, breaking down the original one acceleration and deceleration into two.

[0101] Combination Figure 6 Let's take a look. Figure 6 This is a schematic diagram of an embodiment of the speed of the vehicle after input shaping provided by the present invention. The speed curve of the vehicle after input shaping can be obtained based on the acceleration curve of the vehicle after input shaping.

[0102] Training of anti-sway algorithm for spreader in virtual environment: This invention constructs a quay container gantry crane using Coppeliasim simulation software, and introduces the DDPG reinforcement learning algorithm for training based on the ZV integer input algorithm to initialize the anti-sway strategy. First, a Markov sequence decision model for the anti-sway control of the gantry crane is established, mainly including:

[0103] State s: mainly includes the position and speed of the trolley, the swing angle of the spreader, and the angular velocity of the swing angle of the spreader.

[0104] Action a: The acceleration of the car.

[0105] Reward R: The reward is divided into the reward during the operation of the car and the reward after the car reaches the end of the task.

[0106] During the operation of the trolley:

[0107]

[0108] After the car reaches the destination:

[0109] R(s t ,a t ) = 10*(5-n)

[0110] n represents the number of cycles during which the swing amplitude of the lifting device is less than the preset amplitude.

[0111] In Coppeliasim software, the trolley velocity curve generated by the ZV input shaper is used for simulation, and the data collected during the simulation is input into the offline experience pool database. Taking time t as an example, the trolley position, velocity, spreader swing angle, and spreader swing angle angular velocity at time t are represented by state s. t At time t, the acceleration of the trolley is action a. t After the action is performed, at time t+1, the trolley's position, velocity, spreader swing angle, and spreader swing angle angular velocity are in state s. t+1 The reward received is r t With (s) t a t s t+1 r t The data is stored as a set of data in the experience pool, and the policy network and value network architecture of the DDPG algorithm are established. The DDPG algorithm uses the actor critic algorithm as the basic framework, adopts a deep neural network as an approximation of the policy network and action-value function, and uses the stochastic gradient algorithm to train the parameters in the policy network and value network models.

[0112] Combination Figure 7 Let's take a look. Figure 7 This is a flowchart illustrating an embodiment of the DDPG reinforcement learning algorithm provided by the present invention. The DDPG reinforcement learning algorithm includes the following steps:

[0113] 1. Initialize the policy network and value network (including determining the number of hidden layer nodes, determining the hidden layer activation function and the output layer activation function, and initializing the weights and error values ​​of each node).

[0114] 2. Initialize the experience replay pool and initialize random exploration noise.

[0115] 3. Batch read environment status s t Input into the online policy network and execute action a. t And receive a reward rt and environmental state s t+1 , a set of data (s t a t s t+1 r t The state s is stored in the experience pool R. Simultaneously, the online policy network stores the state s. t+1 Input the target policy network, and the target policy network will adjust the input based on the state s. t+1 Generate the next optimal action a′ t The input is given to the target Q network, and the parameters of the target policy network are directly copied from the online policy network. The online Q network operates based on the state s. t and action a t Calculate the reward function Q(s, a, w) for the action in the current state. The target Q-network calculates the target reward Q′(s′, a′, w′). Update the Q-network using the loss function minimized, and update the policy network using the policy gradient.

[0116] 4. To avoid drastic changes in the calculated target value due to direct updates, which could lead to network oscillations and difficulty in fitting, resulting in bootstrapping, a soft update method is used to update the target strategy network parameters θ′ and the targetQ network parameters μ′, i.e.:

[0117]

[0118] If the experience replay pool is full, the experience replay pool is dynamically adjusted according to the importance of the samples (data with earlier iterations are less important than data with later iterations).

[0119] 5. Repeat steps 3 and 4 until the anti-shake effect meets the requirements.

[0120] Transfer of anti-sway algorithm for real-world environments: Since there are unavoidable deviations between simulation and real-world modeling, how to accurately apply the anti-sway control strategy trained in the simulation environment to the real environment for assembly is a new problem. This invention uses dual-Q network learning to optimize the transfer effect of the strategy.

[0121] Because the policy gradient direction in the actor-critic framework-based DDPG algorithm is a locally maximizing direction, the Q-value of the critic value function network is overestimated. This leads to an artificially high expected reward value for the suboptimal policy in the action policy network, causing a shift in policy network updates. In this case, the TD3 policy utilizes a double Q-network to learn and eliminate overestimation. Based on this idea, this invention transfers the final value function network learned in the simulation environment, which is referred to as the source task critic value function network in the physical prototype experiment. Simultaneously, a randomly initialized target task critic value function network is set up in the physical prototype experiment. To avoid overestimation, the target task critic value function network generally dominates. When the reward calculated by the target task critic value function network is greater than the reward calculated by the source task critic value function network, the source task critic network, the target task critic network, and the actor network are updated.

[0122] Combination Figure 8 Let's take a look. Figure 8 This is a flowchart illustrating an embodiment of the anti-shake algorithm migration from virtual experiments to real environments provided by the present invention.

[0123] The experimental platform for anti-sway physical prototype of quay container gantry crane spreader includes a miniature model of the quay crane, a reinforcement learning anti-sway control system, a load swing angle measuring device, and a trolley position measuring device.

[0124] The angle measuring device measures the swing angle of the suspended load in real time during the trolley's movement, transmitting the swing angle signal to the control system. Similarly, the trolley position measuring device measures the trolley's position in real time during movement, transmitting the position signal to the control system. The trolley position measuring device uses a motor-embedded encoder to extract the trolley position signal and transmits it to the control system. The reinforcement learning anti-sway control system uses the swing angle and trolley position signals as state inputs to the reinforcement learning controller, controlling the trolley's speed and achieving anti-sway control of the gantry crane. The angle measuring device includes a camera, a bracket, and a swing angle measuring host. The camera is mounted on the bottom of the trolley frame via the bracket. The camera acquires images, and the real-time swing angle is measured using deep learning target detection software integrated within the swing angle measuring host. The swing angle signal is then transmitted to the control system.

[0125] The specific steps for anti-sway control of bridge cranes are as follows:

[0126] 1. Initialize all state parameters.

[0127] 2. Enter the target location of the car.

[0128] 3. During the operation of the trolley, the swing angle of the suspended load is detected in real time by the angle measuring device, and the swing angle signal is transmitted to the control system as a status input.

[0129] 4. During the operation of the trolley, the trolley position is detected in real time by the trolley position measuring device, and the trolley position signal is transmitted to the control system as a status input.

[0130] 5. The reinforcement learning control system uses the acquired swing angle signal and the trolley position signal as state inputs to the reinforcement learning controller, and performs real-time control of the trolley's running position based on the output trolley speed.

[0131] This invention addresses the anti-sway task of container spreader on shore. It employs the DDPG reinforcement learning algorithm to train an agent with anti-sway strategies. Using an input shaping algorithm as the initial strategy for the reinforcement learning algorithm controller helps the agent gain a preliminary understanding of the training task, improving sample utilization and algorithm learning efficiency. The inclusion of transfer learning enhances the applicability of the anti-sway algorithm from virtual environments to physical prototype experiments. Compared to traditional open-loop control algorithms like those using integer inputs, the DDPG reinforcement learning anti-sway algorithm achieves good control performance through training without relying on precise modeling. Compared to classic closed-loop controllers like proportional-integral-derivative (PID) controllers, the DDPG reinforcement learning anti-sway algorithm offers higher control accuracy and better adaptability.

[0132] This invention also provides an anti-sway control device for bridge cranes based on deep reinforcement learning, combined with... Figure 9 Let's take a look. Figure 9 This is a schematic diagram of an embodiment of the anti-sway control device for bridge cranes based on deep reinforcement learning provided by the present invention. The anti-sway control device 900 for bridge cranes based on deep reinforcement learning includes:

[0133] Module 901 is used to construct a virtual platform for anti-sway control of bridge and gantry cranes, and to determine the initial strategy of the virtual platform for anti-sway control of bridge and gantry cranes based on an input shaping algorithm.

[0134] The first determining module 902 is used to determine the final strategy of the bridge crane anti-sway control virtual platform based on the initial strategy and the deep deterministic strategy gradient algorithm.

[0135] The second determining module 903 is used to migrate the final strategy to the bridge crane anti-sway control real platform based on the dual-Q network, and determine the anti-sway control strategy of the bridge crane anti-sway control real platform.

[0136] For more specific implementation details of the various modules of the bridge and gantry crane anti-sway control device based on deep reinforcement learning, please refer to the description of the bridge and gantry crane anti-sway control method based on deep reinforcement learning mentioned above. It has similar beneficial effects and will not be repeated here.

[0137] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the deep reinforcement learning-based anti-sway control method for bridge and gantry cranes as described above.

[0138] Generally, computer instructions for implementing the methods of the present invention can be carried on any combination of one or more computer-readable storage media. Non-transitory computer-readable storage media can include any computer-readable medium except for signals themselves that are temporarily propagating.

[0139] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0140] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. In particular, Python, suitable for neural network computation, and platform frameworks based on TensorFlow, PyTorch, etc., can be used. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0141] This invention also provides an electronic device, combined with Figure 10 Let's take a look. Figure 10 This is a schematic diagram of an embodiment of the electronic device provided by the present invention. The electronic device 1000 includes a processor 1001, a memory 1002, and a computer program stored in the memory 1002 and executable on the processor 1001. When the processor 1001 executes the program, it implements the bridge crane anti-sway control method based on deep reinforcement learning as described above.

[0142] In a preferred embodiment, the electronic device 1000 further includes a display 1003 for displaying the processor 1001 executing the deep reinforcement learning-based anti-sway control method for bridge cranes as described above.

[0143] For example, a computer program can be divided into one or more modules / units, one or more of which are stored in memory 1002 and executed by processor 1001 to complete the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in electronic device 1000. For example, the computer program can be divided into the construction module 901, the first determining module 902, and the second determining module 903 in the above embodiments. The specific functions of each module are as described above and will not be repeated here.

[0144] Electronic device 1000 can be a desktop computer, laptop, PDA, or smartphone with an adjustable camera module.

[0145] The processor 1001 may be an integrated circuit chip with signal processing capabilities. The processor 1001 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.

[0146] The memory 1002 may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory 1002 stores programs, and the processor 1001 executes these programs upon receiving execution instructions. The process definition method disclosed in any of the foregoing embodiments of this invention can be applied to the processor 1001, or implemented by the processor 1001.

[0147] The display 1003 can be an LCD screen or an LED screen. For example, the display screen on a mobile phone.

[0148] Understandable, Figure 10 The structure shown is only a schematic diagram of one possible structure of the electronic device 1000. The electronic device 1000 may also include more than one of the following: Figure 10 Show more or fewer components. Figure 10 The components shown can be implemented using hardware, software, or a combination thereof.

[0149] The computer-readable storage medium and electronic device provided in the above embodiments of the present invention can be implemented with reference to the content specifically described in the present invention for the anti-sway control method of bridge and gantry crane based on deep reinforcement learning, and have similar beneficial effects as the anti-sway control method of bridge and gantry crane based on deep reinforcement learning described above, which will not be repeated here.

[0150] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0151] This invention discloses a method and device for anti-sway control of gantry cranes based on deep reinforcement learning. First, a virtual platform for anti-sway control of the gantry crane is constructed, and an input shaping algorithm is used to determine the initial strategy of the virtual platform. Then, a deep deterministic policy gradient algorithm is used to optimize the initial strategy of the virtual platform to obtain the final strategy. Finally, a double-Q network is used to transfer the final strategy of the virtual platform to the real anti-sway control platform of the gantry crane, obtaining the anti-sway control strategy of the real platform. This invention realizes the application of deep reinforcement learning algorithms in the anti-sway control of gantry cranes, improving the performance of anti-sway control of gantry cranes.

[0152] The technical solution of this invention proposes to combine the input shaping algorithm, the DDPG algorithm and the dual-Q network for anti-sway control of bridge and gantry cranes, thereby improving the accuracy of anti-sway control of bridge and gantry cranes.

[0153] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for anti-sway control of a bridge crane based on deep reinforcement learning, characterized in that, include: A virtual platform for anti-sway control of a bridge crane is constructed, and the initial strategy of the virtual platform is determined based on an input shaping algorithm. Based on the initial strategy and the deep deterministic strategy gradient algorithm, the final strategy of the anti-sway control virtual platform for the bridge crane is determined. Based on the dual-Q network, the final strategy is transferred to the bridge crane anti-sway control real platform to determine the anti-sway control strategy of the bridge crane anti-sway control real platform. The initial strategy for determining the anti-sway control virtual platform for the bridge crane based on the input shaping algorithm includes: The initial strategy is determined based on the ZV input shaping algorithm, the natural frequency of the spreader oscillation in the anti-sway control virtual platform of the gantry crane, and the damping ratio of the system in the anti-sway control virtual platform of the gantry crane. The determination of the final strategy for the anti-sway control virtual platform of the gantry crane based on the initial strategy and the deep deterministic strategy gradient algorithm includes: Based on the initial strategy, the state and acceleration of the system in the anti-sway control virtual platform for the bridge crane at any given time are determined, as well as the state of the system in the anti-sway control virtual platform for the bridge crane at the next time after that time. Based on the state and acceleration of the system in the anti-sway control virtual platform for the bridge crane at any given time, the reward of the system in the anti-sway control virtual platform for the bridge crane at any given time is determined. Based on the state, acceleration, and reward of the system in the anti-sway control virtual platform for bridge and gantry cranes at any given moment, and the state of the system in the anti-sway control virtual platform for bridge and gantry cranes at the next moment after that moment, an offline experience base is constructed. Using the offline experience base as an experience replay pool, the final strategy of the anti-sway control virtual platform for the bridge crane is determined based on the deep deterministic policy gradient algorithm.

2. The anti-sway control method for bridge gantry cranes based on deep reinforcement learning according to claim 1, characterized in that, The initial strategy is determined based on the ZV input shaping algorithm, the natural frequency of the spreader oscillation in the anti-sway control virtual platform of the gantry crane, and the damping ratio of the system in the anti-sway control virtual platform of the gantry crane, including: The initial strategy is determined based on the following formula: in, This indicates the amplitude of the first pulse initiated by the anti-sway control virtual platform for the bridge gantry crane. This indicates the amplitude of the second pulse initiated by the anti-sway control virtual platform for the bridge gantry crane. This indicates the trigger time of the first pulse. This indicates the trigger time of the second pulse. This represents the natural frequency of the spreader's oscillation in the virtual platform for anti-sway control of the bridge crane. This represents the damping ratio of the system in the virtual platform for anti-sway control of the bridge crane. As a proportional parameter, the first pulse and the second pulse have the same duration, and the first pulse and the second pulse are used to drive the trolley in the anti-sway control virtual platform of the bridge crane.

3. The anti-sway control method for bridge gantry cranes based on deep reinforcement learning according to claim 1, characterized in that, The determination of the reward of the system in the anti-sway control virtual platform for the bridge crane at any given time, based on the state and acceleration of the system at any given time, includes: When the trolley in the anti-sway control virtual platform for the gantry crane is in operation, the reward of the system in the anti-sway control virtual platform for the gantry crane at any given moment is determined based on the following formula: If the trolley in the anti-sway control virtual platform for the gantry crane reaches the destination, the reward of the system in the anti-sway control virtual platform for the gantry crane at any given moment is determined based on the following formula: in, This represents the reward of the system in the virtual platform for anti-sway control of the bridge crane at any given time. This indicates the state of the system in the anti-sway control virtual platform for the bridge crane at any given time. This represents the acceleration of the system in the virtual platform for anti-sway control of the gantry crane at any given moment. This indicates that the swing amplitude of the lifting device is less than the preset amplitude and the number of cycles.

4. The anti-sway control method for bridge gantry cranes based on deep reinforcement learning according to claim 1, characterized in that, The final strategy, based on a dual-Q network, is migrated to a real-world anti-sway control platform for gantry cranes. The anti-sway control strategy for this platform is determined by: The value function network corresponding to the final strategy is used as the source network of the double-Q network, the target network of the double-Q network is randomly initialized, and the source network is updated. Based on the updated source network, the policy function corresponding to the final policy is updated, and the updated policy function corresponding to the final policy is used as the anti-sway control policy of the bridge crane anti-sway control real platform.

5. The anti-sway control method for bridge gantry cranes based on deep reinforcement learning according to any one of claims 1 to 4, characterized in that, The system status in the virtual platform for anti-sway control of the gantry crane and the system status in the real platform for anti-sway control of the gantry crane include: The position and speed of the trolley, as well as the swing angle and angular velocity of the spreader.

6. A bridge crane anti-sway control device based on deep reinforcement learning, characterized in that, include: A construction module is used to build a virtual platform for anti-sway control of bridge and gantry cranes, and to determine the initial strategy of the virtual platform for anti-sway control of bridge and gantry cranes based on an input shaping algorithm; The first determining module is used to determine the final strategy of the bridge crane anti-sway control virtual platform based on the initial strategy and the deep deterministic strategy gradient algorithm. The second determining module is used to migrate the final strategy to the bridge crane anti-sway control real platform based on the dual-Q network, and determine the anti-sway control strategy of the bridge crane anti-sway control real platform. The initial strategy for determining the anti-sway control virtual platform for the bridge crane based on the input shaping algorithm includes: The initial strategy is determined based on the ZV input shaping algorithm, the natural frequency of the spreader oscillation in the anti-sway control virtual platform of the gantry crane, and the damping ratio of the system in the anti-sway control virtual platform of the gantry crane. The determination of the final strategy for the anti-sway control virtual platform of the gantry crane based on the initial strategy and the deep deterministic strategy gradient algorithm includes: Based on the initial strategy, the state and acceleration of the system in the anti-sway control virtual platform for the bridge crane at any given time are determined, as well as the state of the system in the anti-sway control virtual platform for the bridge crane at the next time after that time. Based on the state and acceleration of the system in the anti-sway control virtual platform for the bridge crane at any given time, the reward of the system in the anti-sway control virtual platform for the bridge crane at any given time is determined. Based on the state, acceleration, and reward of the system in the anti-sway control virtual platform for bridge and gantry cranes at any given moment, and the state of the system in the anti-sway control virtual platform for bridge and gantry cranes at the next moment after that moment, an offline experience base is constructed. Using the offline experience base as an experience replay pool, the final strategy of the anti-sway control virtual platform for the bridge crane is determined based on the deep deterministic policy gradient algorithm.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the anti-sway control method for bridge cranes based on deep reinforcement learning as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the anti-sway control method for bridge cranes based on deep reinforcement learning as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Deep reinforcement learning vibration suppression system and method based on unknown mechanical arm model

    CN114932546A

  • Electronic anti-swing method and system for crane based on Markov decision process

    CN115594083A